ceo @box - your business lives in content. unleash it with AI

Bay Area
Great vision for the future of the creative industry with AI from the one person who can actually predict these things. “As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world. In live action, filmmakers like Steven Spielberg, James Cameron and Peter Jackson embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process. Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process.” Technology has reshaped the creative industry over and over through the years, and has generally always led to an expansion of opportunity or new approaches as a result. And even when the techniques or mediums change, the underlying need for creative skills and taste don’t go away. Just more people are able to leverage those skills effectively when a new technology arrived, and users of the tools discover novel new ways to tell stories.
The World is Changing: AI For Creativity By Jeffrey Katzenberg A few months ago, I sat in my office in Silicon Valley and watched as a tech founder showed me something extraordinary. On the screen was a fully realized, beautifully lit, well-composed animated scene. It was stunning and it made me feel exactly what I felt in 1986 watching Luxo Jr. That was the first time I watched a computer-animated 3D character take a breath and seem, against all reason, to have life. It left me in awe. Later that day, I received a text from an artist I've known for thirty years, 350 miles to the south, in the city where I spent most of my career. After seeing a similar video, she texted: "Is this the end of us?" My answer was, "Certainly not.” I have spent the better part of the last decade in Silicon Valley, but the heart of my career has been in Hollywood. Being deeply connected to both worlds means I have deep loyalties to each and a responsibility to speak honestly to both. In 2023, I said that these new AI tools would cut the time and cost of producing world-class animation by as much as ninety percent within three years. Some colleagues were alarmed, many were furious. There is growing fear and resistance surrounding AI within the creative community. I deeply understand it, because I've spent countless hours walking through animation studios watching gifted artists bent over their desks, rebuilding a single second of film for the tenth time because the ninth version wasn't quite right. I've sat in screening rooms where four years of people's labor played out in minutes, and I knew the name of every person that had spent countless hours bringing those images to life. The creative process is a calling, there's really no other way to describe it. From the outside some see resistance. From the inside, it is love. People do not fight this hard for things they don't care about. The pushback coming out of Hollywood represents the collective effort of people who are deeply passionate about their craft. Is History Repeating Itself? The history here is more complicated than either side may realize. In 1906, the most famous composer in America, John Philip Sousa, published an essay titled “The Menace of Mechanical Music." He warned that the phonograph would become "a substitute for human skill, intelligence and soul." Sousa's fight was not really about the machine, it was about money. The machines were playing his compositions, and the men who built them weren't paying him a cent. His campaign helped create the Copyright Act of 1909. He did not stop the technology. He changed the terms under which it could use his work. A hundred years ago, sound came to the movies. We remember it now as a miracle, and it was. What we forget is who paid for it. Before sound, tens of thousands of musicians made their living in the orchestra pits of movie houses, scoring every film live, every night, in towns all over the world. When the soundtrack arrived, the work of one composer and one orchestra was recorded for a film that went into thousands of theaters. The union fought back with everything it had, taking out newspaper ads across the country warning against the menace of "canned music," one of them showing a mechanical man tearing the strings out of a harp while an angel wept. They were not fools, and they were not Luddites. They were right. Those pit jobs did not come back. And yet (this is the part we have to be brave enough to admit), sound gave us the movie musical, the modern score, sfx, sound design, audio engineering, and an art form vastly larger than the one it disrupted. And it helped keep Hollywood in the forefront of world entertainment for the rest of the century and into the next. The loss was real. And yet the art form expanded. This is a story that has been told over and over again. To resist technology is to risk irrelevance. Just look at Kodak or Blockbuster. To embrace technology is to open doors of new possibility. Just consider Apple and Netflix. What I Learned From Walt Disney In the mid-1980s, I was tapped to lead Disney's animation division at a moment when the studio was at an inflection point. Animation wasn't just another business unit. It was the soul of the company, a medium revered because of Walt's genius and his passion. But the production system was cumbersome and unforgiving. A single movie was 125,000 individual hand-drawn and painted cels, photographed one frame at a time. Every revision carried a cost measured in months. These degrees of difficulty shaped the kinds of stories we could tell. We found our way forward in an unexpected place: Walt himself. The Disney archives held astonishing recordings of Walt explaining his creative process. His own writings. His notes and storyboards. Work product captured at every stage of his process. This was truly a gift. Listening, reading, sitting with the work itself, we heard him talk about character, about emotion, about how an audience feels when a character truly comes alive. He talked about making bold choices and refining a scene until it genuinely moved people. We didn't hear a word about pencils or paintbrushes. In fact, Walt was famous for being a technologist, forever hunting for state-of-the-art tools, often inventing them himself to achieve the images he saw in his head. But he never defined animation by the tools. He defined it by whether the audience believed the character. His principles were timeless. The tools were not. That realization changed everything. We co-developed the Computer Animation Production System (CAPS) with a young Northern California company called Pixar, replacing hand-painted cels with CGI. In The Little Mermaid, the final scene shimmered with a dimensionality and light that the old process simply couldn't achieve. In Beauty and the Beast, the ballroom sequence moved with a cinematic sweep that placed the audience inside the emotion of the moment. In Aladdin, the Cave of Wonders felt vast and alive, and the Magic Carpet became an intricate, compelling character all its own. In The Lion King, the stampede carried a scale and intensity that raised the emotional stakes beyond anything we'd done before. Technology didn't diminish the craft, it expanded the canvas. It gave artists more room to create. A decade later, the canvas expanded again. When Disney released Pixar's Toy Story, it wasn't simply a technical milestone. It was proof that a fully computer-animated film could carry real emotional weight, that it could make audiences laugh, cry, and believe. At DreamWorks, we made the difficult decision to sunset hand-drawn animation and become a fully computer-animated studio. It was the right thing to do, but it was not without pain. It cost talented people their place in an industry where they had worked their whole lives. Some made the leap to the new tools and did the finest work of their careers. Some never did. Tools are never the point. The instruments change with every generation. What endures is taste and imagination. The magical ability to make an audience feel. One of the greatest storytellers of our generation, George Lucas, succinctly captured the eternal essence of this issue: “It’s not the how, it’s the why.” A Distinction With a Difference I asked one of the leading AI models a question that has been challenging me for months. What is the difference between reasoning and creating? Its answer changed how I think about almost everything happening in this industry. It said . . . Reasoning and creating are two distinct cognitive modes, though they also work together. Reasoning is fundamentally evaluative and analytical. It operates on what already exists: facts, premises, evidence. It moves toward a conclusion that was in a sense already implied by the input. Reasoning is constrained by logic and truth. Its goal is to arrive somewhere correct, not to invent somewhere new. Creating is fundamentally generative. It produces something that didn't exist before. And crucially, there's no single right answer waiting to be found. A blank page has infinite valid responses. Creation involves choices that can't be fully justified by logic alone. Taste, intuition and vision fill the gap where deduction runs out. Reasoning is what Silicon Valley has been perfecting. Creating is what Hollywood has been practicing for more than a century. AI today operates almost entirely on the reasoning side of the line. It can deduce, evaluate, optimize, and pattern-match brilliantly. And while it can create, there is a real distinction to being creative. What it doesn’t yet have is those things that make us human: empathy, devotion, serendipity, the kind of creativity that comes from a person trying to say something only they could say. When the bot generates a piece of art, it is not trying to communicate anything. It is statistics, not soul; it is emulating things that have been done. By contrast, human creativity isn’t about repeating patterns of zeros and ones; it is about doing something new. One day, AI may close this gap. Three years ago, the leaders building AI would have called what they are achieving today, improbable, if not impossible. Impossible is no longer improbable. Today, the line between reasoning and creating is real. Even the leading technologists acknowledge we are not there yet. There is no scientific path to crossing this divide that anyone in the field can articulate today. Understanding that gap is where we will find common ground. A Path Forward In 2016, I closed one chapter in Hollywood with the sale of DreamWorks and opened another in Northern California, co-founding WndrCo. We’ve backed more than 50 founders building the next generation of technology and watched how breakthroughs in Silicon Valley emerge, first as experiments, then as platforms, and finally as infrastructure that reshapes entire industries. It's worth remembering that the last great revolution in animation also came from the north. Pixar was a Northern California company, forged not in the conventions of the Hollywood studio system, but in the technological breakthroughs of Silicon Valley. I've spent years on both sides of this bridge. For sure, I don’t have all the answers (take Quibi, for one!). But, from my past and present vantage points of my long career, here is what I see . . . Brilliant people in Northern California building this technology have made something extraordinary. They have earned the right for the rest of us to be, if not believers, at least optimistic that what comes next will be remarkable. But they have not made an artist. The tools are powerful, but they are not what makes a story matter. That knowledge lives 350 miles to the south, inside people whose life's work has informed the very models you are building. The right path forward includes them by design, with credit, with consent, and with compensation. Build this with the storytellers. Not on top of them. Taste is not something that can be synthesized, it is uniquely human. At the same time, Hollywood needs to accept that AI is not going away. The energy they are spending trying to make it disappear is energy they are not spending deciding the terms on which it will exist. And the terms are everything. The north needs something from it that they cannot build and cannot buy: creativity. The kind that takes a blank page and conjures a single right answer where there was none and has held audiences for a century. Without it, the most powerful reasoning engine ever invented will still be missing the only thing that makes a story worth telling. The artists who learn to wield these new instruments will do things the engineers never dreamed of. They always have. Edison invented the motion picture but made terrible movies. It took Chaplin, Lloyd, Keaton and so many others to make movies emotional. Now, the canvas is about to expand yet again. We should decide now that we intend to paint on it. There are so many valuable lessons in history. This has happened many times before, and it was never settled by the technology. It was settled by the terms. Sousa did not stop the phonograph; he helped write the law that made sure composers got paid. And two years ago, when the writers and the actors walked out, they were fighting for the very things Sousa was fighting for in 1906. Consent, compensation, the basic recognition that human creative work has a price that must be paid. The terms of that fight are still being negotiated, but the principle is older than any of us. The tools-versus-no-tools argument is a trap. First, we must all agree that there should be terms. Then we can have the crucial debate about what fairness requires. What I Learned From Steve Jobs Years ago, Steve Jobs said, "It's in Apple's DNA that technology alone is not enough. It's technology married with the liberal arts, married with the humanities, that yields us the result that makes our hearts sing." He was describing a device. But he could just as easily have been describing this tale of two cities. What I See Coming Soon As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world. In live action, filmmakers like Steven Spielberg, James Cameron and Peter Jackson embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process. Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process. Assuredly, I don’t have all the answers, but I am confident that the creative opportunities will expand yet again. How we come through this is a choice. The north has the new tools. The south has the creative soul. The best future will draw on the best of both worlds.
38
13
183
92,345
Muse and Box, together at last
Replying to @hbarra
5/ And some of our favorite productivity connectors, with plenty more in the works: @Box / @github / @meetgranola / @NotionHQ
30
12
404
410,116
What an insane day in AI. The frontier models just became substantially cheaper, with the Opus 5.5 price cuts, and now with GPT-6 Sol and Luna dropping token prices by 50%. The rate at which the cost per task (on a like-for-like basis) drops in AI is unlike any other type of technology in history. And every time the cost of AI drops, the use-cases you can deploy agents against dramatically increase. This is Jevons paradox applied to agents. These improvements will directly lead to broader diffusion of AI in the economy as we can use agents to process all of our data, scan our code for security issues, read through all log data to make decisions, have agent swarms in workflows, and much more. The cost of tokens is directly correlated to these use-cases being opened up at scale.
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
106
84
753
140,493
At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing. Here are some examples of the task wins and performance gains across a variety of industry tests that we performed: • Financial services - due diligence (+39% task accuracy): A year of transaction records, with the job of finding every miscalculation in an acquisition target's pricing tool. Opus 5.5 scored a perfect result on every attempt in half the words Opus 5 used, consuming 82% fewer tokens overall. • Technology - cloud cost analysis (+65% task accuracy): Work out what a company should actually change about its cloud spend. Opus 5.5 picked the right basis for the retention calculation and kept the source data's unit conventions straight all the way through, so the number at the end actually holds up. It took half the time Opus 5 took, with 70% fewer tokens. • Consumer products - client account analysis (+17% task accuracy): Set the onboarding targets for a client account, reading across the signed contract, a satisfaction tracker and a team metrics sheet. The contract never states a senior/junior split, so Opus 5.5 derived it from the 18-person roster and showed the rule it used; several clients had a perfect 10 on individual survey questions, so it averaged each client's responses instead of crowning the single 10. It finished this one in half the time, on 78% fewer tokens. • Clinical diagnostics - data analysis (+15% task accuracy): Malaria rapid-test performance across a dry and a wet season: build the patient records out of two clinical PDFs, compute positive test rates by season and gender, and test whether parasite counts really differ between test-positive and test-negative patients. Opus 5.5 caught that the two groups' standard deviations differed more than 100-fold, re-ran it the right way, and found the dry-season difference didn't hold up after all. This accuracy gain came with a final answer that was half the length of Opus 5's, and also needed 78% fewer tokens end to end. Customers will be able to build AI Agents with Opus 5.5 shortly in the Box AI Studio.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
59
28
384
148,025
The monetization potential of personal agents that are transacting on your behalf you is quite significant. If you imagine agents that are perfectly capable of handling an arbitrarily complex task end to end, then eventually a substantial amount of commerce inevitably goes through them. You’ll start by slinging your daily simple and annoying tasks at the agent. Then as people get used to it, they’ll just start to throw more complex tasks at the agent, ultimately leading to even more spend through these systems than what they were doing before. If you can bring down the friction for commerce and services, then you end up spending even more. This thus creates a ton of opportunity for the agent providers (Muse, etc.), but also completely new opportunities to build the layer that the agents want to interact with (commerce, local, b2b services, etc.). Win/win for multiple layers.
It’s funny, Meta went from having my Instagram and WhatsApp data to now having access to my email, calendar, DoorDash, Amazon and pretty much everything. In the last 24 hours, it bought me socks, ordered my Whole Foods groceries, booked a cleaning service and got me a burger for dinner. Meta’s last disclosed North American Facebook ARPU was around $227/year, largely from ads. I suspect it can push that number significantly higher now that it understands not only what I look at, but what I need, what I buy and what I’m planning to do. Also the much bigger opportunity might be becoming the aggregation layer between me and the entire internet. If Meta can take even a tiny percentage of the commerce it facilitates, or of the money it saves me, this could become enormous! It already saved me $200 by canceling subscriptions and services I no longer needed. This feels much bigger than better ad targeting. Ads are useful but giving me money back is better imo. One additional thought: the agent is increasingly making the decisions for me. I knew nothing about that burger place. The agent researched it, told me which burger I should order, and I just said “okay” without giving it much more thought. Agents are becoming the decision makers in both B2C and B2B. Increasingly, every business will be selling not just to humans, but to their agents. Everything becomes B2A: business to agents.
51
21
237
104,230
AI agents will use software 100X more than people ever did. Even as interfaces begin to fade into the background as you primarily interact with agents, those agents still need many of the core primitives that people have used. In fact, in many ways these core primitives become even more important when agents can take destructive actions in our systems or where the context they’re accessing with make or break the workflow. This will be true of where you house your CRM or ERP or structured or unstructured data platforms. The platforms that can best act as the security layer and guardrails for agents, manage the data for agents and people, and orchestrate the business logic for workflows have a huge opportunity right now. This is true for brand new startups as well as existing platforms that can move fast enough.
Box CEO @levie says AI has been “unequivocally” a net positive for software. “You look at infrastructure providers—the Cloudflares of the world—totally on fire, because agents need sandboxes, they need compute, they need network, they need gateways. Great business.” “For us, we have reaccelerated growth far past our internal plans because it turns out that enterprises need core systems to be able to manage their unstructured data.” “Agents need to be able to work with that unstructured data to make decisions or move information through a workflow—whether it’s all of your contracts, your research materials, marketing assets, or financial documents.” “Now imagine an enterprise with 10,000 employees. How do you ensure that those agents continue to go after the actual canonical records, the actual documents that are the real sources of truth and the authoritative versions of that data?” “That means you need platforms.” “And all of that is creating more value for these platforms.”
78
30
315
108,850
Literally impenetrable from agent swarms
IT security in 1990s
167
228
6,381
242,755
Personal agents are the ultimate manifestation of “build something that agents want”. The form factor of a product like Muse is you want to be able to hand off a task to the agent and ensure that it is fully completed end to end. To do this, the agent must be able to successfully operate with your tools or use its own to complete the task. Use your MCP or CLI, easily navigate your site, be able to transact, and more. The new attention you need to compete for is not from the user itself but instead for the agent. This means that the tools that allow agents to order food, handle ecommerce transactions, book flights, work with the local economy, and interact with our data and information best, are the ones that will get used the most. This will ultimately be the biggest opportunity and shakeup in consumer tech since the App Store itself.
zuck essentially launched a new app store for agents today and Meta’s perfectly placed to crush it, you’re looking at a multi-trillion dollar opp if they pull it off: - instead of apps, agents get equipped with connectors i.e. plugins to any app, service or software tool. - suddenly your agent goes from a useless chatbot to an action-oriented helper that gets shit done for you but it gets even better - meta has collected ALL the data about you. they know what you want, when you want it which means they know exactly WHICH connectors to list in a marketplace and HOW MUCH to charge because they know customers will pay. it’s a masterclass. muse isn’t just some generic agent, it’s sitting on a treasure trove of data and meta can build a trillion dollar business on top of it.
95
62
602
157,752
Jev will be super helpful for agents to make split second decisions in workflows, data classification, judgment calls, and hundreds of other use-cases in the enterprise. Here's a quick demo with Box and Jev to make that real. The demo pulls an incident report from Box, asks whether it's customer-facing and how severe it is, moves the file into escalate, monitor, or review folders, and sets a metadata template instance with the result. This all happens nearly instantly and at almost no cost. You can imagine this in insurance claims, contract management, loan processing, security reviews, customer log analysis, and so on. Definitely a great new class of AI use-case.
74
76
669
114,828
Agents already make up the majority of inference. This will quickly trend toward nearly all inference over the next year or two. The vast majority of tokens used in the world will be agents that are executing unbelievable amounts of tasks for us in the background 24/7. Agents will be deployed to read all code changes to secure our software, process all of our data inside of workflows, handle a significant majority of the research that goes into recruiting and customer prospecting, review every event stream and log from every system in the world, go out and execute tasks for us in our personal lives, and hundreds of other use cases. The rate of new agents coming online that will be consuming insane amounts of tokens is not slowing down. Just in the past week I’ve introduced multiple completely new workflows that would not have been possible technically even a month ago. Incredible time to be doing anything in inference and of course building on top of all of this.
AGENTIC TRAFFIC NOW MAKES UP MORE THAN 70% OF ALL INFERENCE TRAFFIC 🚀 Agentic workloads are characterized by four elements: 🟠 Multi-turn: a session includes tens or hundreds of turns, leading to high potential KV-cache reuse. 🟠 Long context: system prompts, tool definitions, and the large number of turns make context accumulate quickly. 🟠 High prefix reuse: since the conversation progresses linearly, where output from turn n-1 is concatenated to turn n (typically), most context can be served from KV cache rather than recomputed (this depends on the amount of storage available to store KV tensors). As n grows, the ratio of cached input relative to uncached input typically tends towards 1. 🟠 Sub-agent bursts: a session launches multiple short-lived sub-agents with fresh context, which create bursty KV-cache patterns.
62
33
282
87,461
Incredibly exciting that there are entire universes of AI innovation that still exist that weren’t even on most of our radars. Being able to process information insanely quickly, at crazy low costs, with high levels of capability is huge for a wide number of enterprise tasks. Data classification tasks, routing decisions inside of a workflow, decision making when handed a particular domain problem, quick judgment calls about safety or security, and more all are the gates in a large number of processes. This model and approach could be quite cool in agentic workflows in the enterprise.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
64
39
529
112,899
Aaron Levie retweeted
Claude’s new Slides experience can turn deal materials in Box into an editable credit committee briefing. In our demo, @claudeai reviews a borrower’s deal workspace, creates a five-slide presentation with financial visuals and source references, and flags conflicting information for review. We leave a comment directly on a slide, Claude revises it, and the finished deck is exported as PowerPoint and uploaded back to Box through MCP for the credit team to review. Connect Box to Claude to create and refine presentations using your enterprise content.
Replying to @claudeai
You can also now make decks, docs, and designs in your conversation. Draft the one-pager in Claude Docs, turn it into a deck with Claude Slides, and mock up a matching visual in Claude Design, all from one place.
4
3
37
735,174
There’s a massive chasm between the power of AI models and the ultimate workflows that enterprises are trying to automate. This gap is the opportunity for the applied AI layer to fill. You need to connect the intelligence to workflows, often reengineer processes, aggregate the right context and data, allow for the right human in the loop experiences, drive change management, do domain specific evals, manage the security and governance of the data and process, and much more. We’re going to see this layer emerge in every vertical and horizontal category. And ironically, even as models improve at incredible rates, this layer still must exist - and may become even more important and useful. Greater capability enables even more complex tasks to be tackled, amplifying the challenges if you don’t do this well. Was super fun chatting with @sonyatweetybird on all the things going into AI diffusion.
Well.. @levie and I filmed this episode of Training Data a week or two ago, when the “current thing” was Doug Leone’s novacaine root canals instead of pacing the frontier… Simpler times! But Aaron’s advice on reinventing yourself and your company for AI is timeless. Aaron founded @Box 20 years ago. It sits on hundreds of billions of enterprise files, and he's bet the company on agents that can read every one of them. He's also one of the most wired-in people in AI, on every cap table and, by his own admission, 95% Twitter-educated. He’s the rare CEO who can straddle both the internet AND has the ear of CIOs. His core argument: (1) the gap between what a model can do and what an enterprise workflow actually needs is vast, and closing it is a lot of software; (2) diffusion of AI outside of coding will take far longer than Silicon Valley thinks, and that slowness is exactly where the applied layer's value comes from. The conversation covers: — why application companies are the hottest neolabs, and why the LLM-wrapper thesis is finally working — the fox-guarding-the-henhouse problem with letting model providers route your tokens — work slop, and why we accept AI-written code but flinch at AI-written decks — how Box built its agentic harness and why it beats raw API access on accuracy and latency — the open-weights paradox: closed labs and open models both growing exponentially at once — what continual learning has to solve before it works for a lawyer with five matters and a Chinese wall — why 90% of enterprise tokens in five years will come from tasks no human kicked off — the mandate for founders right now: whoever gets it to the customer wins 0:00 – Introduction 1:55 – Are application companies the hottest neolabs? 6:56 – Will the labs move up the stack? 12:34 – Box and betting the company on AI 16:50 – Hero use cases: reading a million contracts and long-running agents 18:42 – Work slop: why AI code is embraced but AI content isn't 24:08 – Building Box's agentic harness and the evals that matter 27:23 – The state of the model race 29:25 – Open-weight model adoption in the enterprise 32:34 – Memory, continual learning, and what belongs in the weights 37:29 – Box Labs and systems of record in a world of agents 44:55 – Will chat be the dominant UI for enterprise AI? 48:00 – Why coding diffused fast and the rest of knowledge work hasn't 54:31 – Staying wired in, making a company AI-first, and what it takes to win
86
51
358
119,339
Aaron Levie retweeted
Well.. @levie and I filmed this episode of Training Data a week or two ago, when the “current thing” was Doug Leone’s novacaine root canals instead of pacing the frontier… Simpler times! But Aaron’s advice on reinventing yourself and your company for AI is timeless. Aaron founded @Box 20 years ago. It sits on hundreds of billions of enterprise files, and he's bet the company on agents that can read every one of them. He's also one of the most wired-in people in AI, on every cap table and, by his own admission, 95% Twitter-educated. He’s the rare CEO who can straddle both the internet AND has the ear of CIOs. His core argument: (1) the gap between what a model can do and what an enterprise workflow actually needs is vast, and closing it is a lot of software; (2) diffusion of AI outside of coding will take far longer than Silicon Valley thinks, and that slowness is exactly where the applied layer's value comes from. The conversation covers: — why application companies are the hottest neolabs, and why the LLM-wrapper thesis is finally working — the fox-guarding-the-henhouse problem with letting model providers route your tokens — work slop, and why we accept AI-written code but flinch at AI-written decks — how Box built its agentic harness and why it beats raw API access on accuracy and latency — the open-weights paradox: closed labs and open models both growing exponentially at once — what continual learning has to solve before it works for a lawyer with five matters and a Chinese wall — why 90% of enterprise tokens in five years will come from tasks no human kicked off — the mandate for founders right now: whoever gets it to the customer wins 0:00 – Introduction 1:55 – Are application companies the hottest neolabs? 6:56 – Will the labs move up the stack? 12:34 – Box and betting the company on AI 16:50 – Hero use cases: reading a million contracts and long-running agents 18:42 – Work slop: why AI code is embraced but AI content isn't 24:08 – Building Box's agentic harness and the evals that matter 27:23 – The state of the model race 29:25 – Open-weight model adoption in the enterprise 32:34 – Memory, continual learning, and what belongs in the weights 37:29 – Box Labs and systems of record in a world of agents 44:55 – Will chat be the dominant UI for enterprise AI? 48:00 – Why coding diffused fast and the rest of knowledge work hasn't 54:31 – Staying wired in, making a company AI-first, and what it takes to win
9
3
80
390,490
Aaron Levie retweeted
We have raised $200M at a $5B valuation to scale self-improving software development in the enterprise. @FactoryAI has grown to serve hundreds of thousands of developers at companies including RBC, Adobe, Nvidia, T-Mobile, and Palo Alto Networks. We will use this capital to accelerate our investments in research, product, and global go-to-market.
161
68
780
274,688
We probably need to all update our sense of what’s coming in terms of agentic workloads with the combination of agent swarms, better computer use, the next wave of APIs and MCPs coming online, and new form factors like Muse or Instinct, vertical enterprise agents, background workflow agents, and other products that are emerging right now. We’re going to throw agents at vastly more tasks in our professional and personal lives than we had initially imagined. The kind of information that agents will go out and find for us and do work for us in the background is going to be 100X more volume than what we can imagine previously prompting in a single session. You’re going to have agents go out and surgically recruit talent for you 24/7, look for every signal in your customer’s business for when to pitch them, process every single transcript and conversation for product insights, review every line of code written for security issues and bugs, brute force test your systems for issues, and 100s of other tasks. We’re probably 1% of the way into the shape of what all these agents look like, where they get deployed from, how they get managed, how we budget for them, and so on. But it’s inevitably going to happen at a scale that wouldn’t have been fathomable before.
155
96
822
119,982
Protecting enterprise data in a world where agents are using our systems 100X more than people ever did is going to be one of the more complex security and governance challenges of the 21st century. Importantly, security and productivity gains are inexorably linked in the world of AI. If you give an agent too much unfettered information access, it will be difficult to truly control and protect your data; and conversely, if you lock everything down completely, you won’t get any real productivity gains from AI. We need all new ways to protect our systems, environments, and structured and unstructured dada in the enterprise in an intelligent way by modernizing our approach to security and governance. At Box, as one example, we’re building new intelligent ways to protect enterprise data and agentic use of that information. A recent update in Box Shield is to provide granular controls on what content agents can and can’t work with based on document classification level. We’re also working on other features that can automatically detect and alert (or block) when data is being accessed or used in unusual or anomalous ways by agents. And this is just the start. There’s a ton more innovation coming across the entire industry - from the labs like OpenAI or Anthropic; security platforms like Palo Alto Networks, Cisco, CrowdStrike, Okta; or startups like Enclave, Method, Alterion, Runlayer, and many many others - to rethink how we protect information in the world of AI agents. Exciting and wild times ahead.
AI tools will keep changing, but content policy needs to travel with the content. See how Box Shield applies classification-based access controls to what AI agents can actually read, not just what they can download. 👇
48
21
112
298,045
“Pacing” can be somewhat of a trigger word because it sounds like an arbitrary slow down of capability or a way to hobble competitors through undue regulation. However, the specific improvement goals laid out in Dario’s are absolute necessities in AI development. We expect equivalent in areas like aerospace, life sciences, health care, and other industries, and AI development likely shouldn’t be that different. AI is going to be the technology underpinning our financial trading systems, our medical devices, our biotech breakthroughs, defense systems, government workflows, and many more mission critical domains. So it completely stands to reason that we want these systems to be safe and “aligned”. How we get there -without meaningfully slowing down innovation progress or reducing competition- is one of the most complex questions in the 21st century, but the need is clearly real.
78
34
210
52,956
Good post. Don’t agree with all of it, but it represents many of the actual practical realities for the general path forward in frontier AI whether we like it or not. At the level of capability we’re seeing from models there will inevitably be some form of coordinated self-regulation of the industry; that’s generally a good thing. The big question will be if everyone can agree to what that is. However, the way things are headed politically there’s a chance that the labs won’t even get a say in it. And then the bigger question yet is do all the countries participate or not. Any “slow down” hinges on broad participation, which from a game theory standpoint isn’t likely until the risks are more severe and obvious. Expect this to be pretty messy for a while.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
69
25
278
109,738
This one is very cool. Now you can mount Box to agent sandboxes to make it far easier for an agent to read and write files on the agent's computer. As AI agents execute critical workflows in the enterprise, they're going to need the same primitives that people have had.
🆕 @OpenAIDevs Agents API gives developers a hosted sandbox where agents run tools, coordinate work, and handle complex tasks. Box Mount brings enterprise content directly into that sandbox as files the agent can read, reason across, and produce new work from, using normal shell commands and file paths while Box Mount handles two-way sync automatically. In our demo, a lead agent mounts a deal room from Box, reads five source files, launches three specialist agents in parallel, and writes four reports back through Box Mount with an initial verdict. When a customer revises the MSA in Box, Box Mount syncs the new version into the same sandbox and the agents reassess automatically, refreshing all four reports while preserving the revised MSA as version two in Box. The result is a shared workspace where people and agents work on the same governed Box content without building custom file-transfer logic, with permissions, versions, and audit history preserved throughout. Box Mount is in private preview. Watch👇
46
22
116
83,255