Founder @attaapp. ML PhD. Previously worked on FaceID at Apple.

San Francisco, CA
Pinned Tweet
I’m excited to share that @attaapp has been acquired by @Replit ! I started Atta with my co-founder, @omarshaik , because we thought AI would change how people explore and build with data. I wrote the first lines of the canvas POC for Fastview, which later became Atta, in @Replit. Now Atta’s charts and reports are live in @Replit . It’s been great working with @pirroh and the team, and I’m excited to be part of what @amasad and the @Replit team are building. A special thank you to @mejoff and @shawnvc and the @MenloVentures Labs team for believing in us and supporting us along the way. I’m incredibly grateful to our team, customers, investors, advisors, friends, family, and everyone who believed in us early or gave us feedback along the way. Read more here: replit.com/blog/replit-acqui… ❤️
3
3
28
7,132
abk retweeted
Big news! Atta is now part of @Replit. Now, I can uncover business insights and build interactive in-line charts directly within Replit. Just start with a prompt and let us handle the rest. This channel will no longer be active after 09/30, but continue following our journey over at @Replit. To new beginnings.
3
4
9
1,923
Try @attaapp charts in @Replit today!
Data visualization now lives directly in Replit. We’ve acquired @Attaapp to bring their high-quality, purpose-built charting into conversations. So you get interactive, in-line charts for your next board meeting or sales pitch. No SQL. No rebuilding over and over again. Just ask Replit and get analysis you can explore, with charts you’re ready to share.
2
85
abk retweeted
Data visualization now lives directly in Replit. We’ve acquired @Attaapp to bring their high-quality, purpose-built charting into conversations. So you get interactive, in-line charts for your next board meeting or sales pitch. No SQL. No rebuilding over and over again. Just ask Replit and get analysis you can explore, with charts you’re ready to share.
8
7
83
7,590
Adaptive Depth RNN with 32 dim hidden state with Jev 😆
Introducing Looped Jev. System 1 thinking at System 2 price.
1
4
247
abk retweeted
the right way to use model capabilities is not to ship 10x more features to prod it's to spend more time understanding your users, trying experiments, building prototypes, learning about things you don't understand so that you can ship things that actually work
275
615
8,210
381,394
What's even a "a refreshing form factor"
guys Jev is not a “universal classifier” 😭 what even is that it’s - a very strong model - where a great team made tradeoffs it can classify well - by creating very high quality data - that aligned with the domains/behaviors they wanted the model to be good at - in a refreshing form factor for intelligence ask the model to play chess, it’s all tradeoffs you don’t need a general purpose classifier to have an incredibly useful model
1
1
339
Is Jev too dangerous to release?
1
52
+1
Prediction We’re gonna have Continual Learning adoption-ready in Local AI by the end of 2027 Markdown harnesses are bloat, in-context learning doesn’t work, Continual Learning is the only way forward
43
T-800
45
A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer". It's always interesting to read about new or different approaches (including rumors about what the closed labs may be up to), but let's debunk this a bit. About 2 months ago, I shared the architecture details of Nanbeige, for example, where "Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters." Yes, that's it. The looped transformer idea is just reusing layers in the transformer block. In the case of Nanbeige, the main idea is to reuse the same 22-layer stack (=transformer block) twice instead of once. So, effectively it extends the 22-layer architecture to 44 layers, but without duplicating the weights. In simple terms, this roughly doubles the size of the model (if we ignore the embedding and output layers for a second). But instead of requiring 2x the storage and RAM to host this model, it stays at the same size since we reuse the components. However, it's almost 2x as expensive in terms of compute, because we run the embedded text through almost 2x as many layers. Why? In the Nanbeige 4.2 technical report, the researchers found that two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. (More passes gave barely any gains but made the training much slower and much more expensive.) While, as far as I know, Nanbeige 4.2 is the first notable open-weight model that adopted this approach, the idea goes back to the NeurIPS paper "Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation". Actually, this paper proposes a mechanism that is a bit more sophisticated by adding a learned router that determines whether each token receives one, two, or more passes. So, easy tokens can exit early while harder tokens receive additional computation. In sum, Astra may be a really good model, but this shouldn't be about this "looped transformer aspect," which is just a tiny architectural tweak. Also, the statement "the new technique works in a way that obscures some or all of the AI's reasoning, otherwise known as 'chain-of-thought'" is not necessarily true with respect to the looped transformer method. It's possible that The Information journalist refers to some other technique or misunderstood the looped transformer method. Reusing layers does not by itself suppress visible chain of thought. It adds computation in hidden states before the next token is emitted, just as ordinary transformer layers do. But based on the information we have, the only plausible interpretation here is that if a model uses more of these recurrent passes, it may need to generate fewer intermediate reasoning tokens. So then more of its computation happens in latent activations that cannot be read as text. But we would get the same effect if we were scaling up the model size, like GPT 5.6 Luna -> GPT 5.6 Sol.
139
588
4,108
386,976
abk retweeted
Agents made software cheaper but made coding expensive. Today, together with @OpenAI, we’re changing this:
Replit Free Mode, powered by @OpenAI GPT-5.6 Luna. Let’s make intelligence accessible to everyone.
193
253
4,068
605,607
Nice work!
Introducing Atta Desktop. Business analysis built for AI-pilled teams using Claude Code, Codex, etc. Claude + a harness of semantic definitions, skills, knowledge files works great to answer quick data questions. But sometimes you need an expansive space to think. You need more polished outputs than whatever the random token generator comes back with. That's where Atta Desktop comes in. Pull your data from Snowflake, run formulas in Excel, lay out scenarios in Atta, and download a gorgeous chart to drop into a PowerPoint, all directly from Claude Code. Data stays on your machine and in your corporate agent environment. If that sounds like your setup, download Atta at the link below and let me know what you think!
5
1
8
2,555
abk retweeted
Download Atta Desktop: atta.app/downloads
Introducing Atta Desktop. Business analysis built for AI-pilled teams using Claude Code, Codex, etc. Claude + a harness of semantic definitions, skills, knowledge files works great to answer quick data questions. But sometimes you need an expansive space to think. You need more polished outputs than whatever the random token generator comes back with. That's where Atta Desktop comes in. Pull your data from Snowflake, run formulas in Excel, lay out scenarios in Atta, and download a gorgeous chart to drop into a PowerPoint, all directly from Claude Code. Data stays on your machine and in your corporate agent environment. If that sounds like your setup, download Atta at the link below and let me know what you think!
1
2
982
Today, we’re launching Atta Desktop. Agent harnesses are becoming the new operating environment for knowledge work. They can access your code, files, data warehouse, semantic layer, and systems of record. But data analysis is inherently visual, spatial, iterative, and comparative, yet agents still do most of their work inside a one-dimensional conversation. Data analysis needs somewhere to unfold: charts placed side by side, hypotheses explored in parallel, conclusions annotated, and results organized into a coherent story. Atta Desktop gives your agent that missing visual surface.
2
2
27
45,134
It is a local, infinite canvas that connects to Claude Code, Codex, and other compatible agents through MCP. You bring the harness you already use, along with its context, tools, and access to your data. Atta gives it a place to explore visually, create professional charts, and turn analysis into executive-ready work. And because Atta Desktop is local-first, your analytical data remains on your machine. Atta only connects externally for account management.
1
1
196
Every capable agent will need surfaces designed for the work it performs. For data analysis, we believe that the surface is Atta. atta.app/downloads
158
Now more than ever.
45
207
2,946
197,971
Many people think any given ML project is 99% training. In reality, it’s 50% evaluation, 40% data cleaning, 8% integration, and 2% training. The first two set the noise floor for learning. No ML magic matters; the model cannot lower the noise floor, as that’s the optimal bound of Shannon encoding of your data. Thus, not a single day goes by without me thinking about ontology. Even the old labels have to be constantly reviewed.
546
1,270
11,013
17,876,456
Don't get fooled by token prices!
4
88