Very realistic overview. Hard to discern hype from reality about Jev so here are some bitter facts:
- Jev can’t write text: it’s not an autoregressive model, so forget about anything related to coding or writing
- Jev can’t see (yet): text only, althought there is nothing fundamentally preventing it from being multimodal. I’m sure they will add support for images?
- Jev can ONLY choose a set of predefined actions: the good thing about computer and browser use is that by very nature it’s a STATE -> ACTION model. The problem with long running state action models is that the reasoning and state understanding becomes extremely important.
- Jev is extremely cheap and fast
It opened up my eyes into what’s possible. We are often stuck in optimizing problems inside the box. This is one the real “think outside the box” solutions to problems we have been trying to solve.
I am extremely hyped about the future of computer and browser use. Latency matters, and people are clearly hyped about it.
We can surely combine some sort of global state understanding (LLM) with super fast actor model (System One). Obviously it’s possible - FSD and robotics companies have solved this already.
How hard can it be to apply the same thing to browser use?
Jev is awesome but for the love of god please STOP posting fake demos