What is real
I have been running Hermes Agent, the Nous Research one. Terminal, messaging, a box that stays on. It can write or refine skills after a messy task. You feel that the second time you ask for the same thing.
That part is not vapor. Memory stores the facts; skills store the procedure. The practical result is less re-explaining and fewer repeated mistakes.
Where it wanders
It is also not a company. It is a sharp tool with a learning loop, and it will happily wander if you do not point it. Fine for nights. Not how I would bill a client on Monday.
The product work is everything around the model: permissions, stop conditions, source quality, task ownership, retries, and proof that the requested thing actually happened.
What survives
My current take: memory and skills often beat a slightly longer context window. The agent that remembers how you file invoices is more useful than the one that scores two points higher on a public benchmark.
I will keep breaking it. When something survives a week of real use, I will write that down here.