Aaron Levie
We’re already starting to see what kind of new jobs AI is creating. AI requires significant technical work and surrounding services to deploy into the economy. This means jobs for AI engineers that build applied AI products sold to (or within) enterprises, FDEs to deploy agents into companies, new services firms for deploying AI, and more. And even the published stats of new jobs undercount all the existing jobs that are transitioning to new areas of AI work in an enterprise. Many prior data, research, and software jobs large enterprises are also being repositioned for working with AI. Every bank, life sciences company, manufacturer, and even law firm is bringing on more technical talent -or repositioning existing roles- to help with agent deployment in their companies. It’s a lot easier to picture what AI can replace vs. what it creates until it starts happening. Now we’re seeing what this looks like.
中文: 我们已经开始看到人工智能正在创造什么样的新工作岗位。 人工智能需要大量技术工作和周边服务才能向经济领域投入。这意味着人工智能工程师能够将应用人工智能产品出售给(或内部)企业,FDE用于向企业部署代理,以及为部署人工智能而开发新的服务公司,以及更多工作。 甚至已公布的新工作岗位统计数据也低估了企业中所有正在向人工智能工作新领域过渡的现有工作岗位。许多以往的数据、研究和软件工作大型企业也正在被重新定位以用于人工智能。 每家银行、生命科学公司、制造商甚至律师事务所都在引进更多技术人才,或重新调整现有职位,以帮助其在公司中部署代理人员。在人工智能开始发生之前,要了解哪些能取代它所能替代的,这要容易得多。现在我们看到这看起来是什么样子。
Aaron Levie
This metric may tell us more about the business model of consumer AI than the state of AI. AI for consumers almost inevitably will be subsidized and monetized via commerce, ads, or purchases of devices and other services. This metric may level off lower than we think.
中文: 这个指标可能告诉我们更多关于消费者人工智能的商业模式,而不是人工智能的现状。面向消费者的人工智能几乎不可避免地会通过商业、广告或购买设备和其他服务获得补贴和变现。这一指标可能比我们想象的要低。
Aaron Levie
AI agent adoption is still very bimodal right now. You have coding and coding adjacent tasks which have taken off, and then everything else. And even within coding, you have a very wide continuum of adoption patterns, with some likely a small percentage of developers deploying background agents working on projects in parallel, and the rest of everyone else still working with agents 1:1. Then, there’s the entire rest of knowledge work where real agentic adoption is still very early. The reason for this is that most of the workflows still need to reengineered to work with agents. This isn’t the same as deploying a chat system for better research efficiency, but instead requires workflows to be rebuilt, data to be wired up in new ways, new practices for accountability and liability of decisions when agents are in the mix, governance and compliance paradigm changes, security upgrades, and more. We’re still so unbelievably early in what this is going to look like outside of a few categories of work right now. This is why you can basically expect 100X more agent adoption from what we’ve seen so far.
中文: 人工智能代理的采用目前仍然非常二模态。你有编码和编码相邻的任务,这些任务已经起飞,然后是其他所有任务。 即使在编程中,你也拥有非常广泛的采用模式,其中可能有一小部分开发人员会同时部署后台代理来处理项目,而其他人仍然以1:1与代理合作。 然后,在整个知识工作中,真正的特化采用仍然非常早。其原因是大多数工作流程仍需重新设计才能与代理人员合作。 这与部署聊天系统以提高研究效率不同,而是需要重建工作流程,以新的方式连接数据,在代理人员处于混合模式、治理与合规范式、安全升级等因素中,为责任追究和责任决策提供新的实践。 目前,在少数几类工作之外,我们仍处于如此令人难以置信的初期。这就是为什么从我们目前所看到的内容中,基本上可以期待100倍的代理采用率。
Aaron Levie
RT @talhof8: At @EnclaveAI, we're building the platform to protect the enterprise: an always-on security program that finds real risk, prepares the fix, and verifies the attack no longer works. Figuring out which models are best for cybersecurity is an increasingly difficult task. Open-weight models do different types of things at different levels. Some are very good at vulnerability discovery, others at exploitation, others at triage or remediation. And a new one ships every few weeks. We're announcing Enclave Router. Enclave Router allows you to deploy cyber inference for your codebase, across any harness that you want. One OpenAI-compatible API for the best cyber-capable open-weight models, benchmarked per security task, with zero data retention. We're giving you $100 free to start using it, just reply here or DM me. https://router.enclave.ai/
中文: RT @talhof8:在@EnclaveAI,我们正在构建保护企业的平台:一个始终持续安全的程序,能够发现真实风险,做好修复准备,并验证攻击不再有效。 弄清楚哪些模型最适合网络安全,是一项越来越困难的任务。开放式重量模型在不同层面上进行不同类型的事情。有些擅长发现漏洞,有些则擅长利用,另一些则擅长分诊或补救。每隔几周就有一艘新船。 我们正在宣布《飞地路由器:飞地路由器”。Enclave Router 可让您在任何需要的线束上为代码库部署网络推理。一个兼容OpenAI的API,用于实现最佳的网络可控开放权重模型,可对每个安全任务进行基准测试,但数据保留率为零。 我们免费给你100美元开始使用,只需回复这里或我。
🎬
视频
Aaron Levie
A big trend with most enterprises that I talk to is deploying internal FDEs into departments in their organization to help bridge the capabilities of AI into their underlying workflows. This requires a high level of technical expertise, an understanding of AI, and the ability to understand the workflows and processes that the enterprise is trying to automate. There’s no shortcut to getting automation without this combination of skills (either in one person or multiple). What’s exciting is this is an entirely new function in most enterprises, which is going to create a ton of new roles in the economy. “Every generation of technology creates jobs that didn't exist before it, jobs that arise from an immediate need in an emergent industry. The Automation Engineer is a key example of those jobs for this generation. Nobody has ten years of experience doing this yet, because ten years ago the tools that made it possible didn't exist (it didn’t even exist two years ago).” If you have software skills and are diving into AI, this is an area to go deep on.
中文: 与大多数企业合作的一个大趋势是,将内部FDE部署到其组织中的部门,以帮助将人工智能的能力融入其底层工作流程中。 这需要高度的技术专长、对人工智能的理解,以及理解企业试图自动化的工作流程和流程的能力。没有这种技能组合(无论是一个人还是多人),实现自动化就没有捷径。 令人兴奋的是,这在大多数企业中是一项全新的功能,将在经济中创造一大新角色。 每一代技术都创造了前所未有的就业机会,这些就业机会源于新兴行业的迫切需求。自动化工程师是这一代人从事这些工作的关键范例。目前还没有人有十年的工作经验,因为十年前制造这种可能的工具并不存在(甚至两年前也不存在)。 如果你具备软件技能并深入人工智能领域,这是一个需要深入探讨的领域。
Aaron Levie
A big trend with most enterprises that I talk with is deploying internal FDEs into departments in their organization to help bridge the capabilities of AI into the underlying workflows in the organization. This requires a high level of technical expertise, an understanding of AI, and the ability to understand the workflows and processes that the enterprise is trying to automate. There’s no shortcut to getting automation without this combination of skills (either in one person or multiple). What’s exciting is this is an entirely new function in most enterprises, which is going to create a ton of new roles in the economy. “Every generation of technology creates jobs that didn't exist before it, jobs that arise from an immediate need in an emergent industry. The Automation Engineer is a key example of those jobs for this generation. Nobody has ten years of experience doing this yet, because ten years ago the tools that made it possible didn't exist (it didn’t even exist two years ago).” If you have software skills and are diving into AI, this is an area to go deep on.
中文: 与大多数企业合作的一个大趋势是,将内部FDE部署到其组织中的部门,以帮助将人工智能的能力融入组织的底层工作流程中。 这需要高度的技术专长、对人工智能的理解,以及理解企业试图自动化的工作流程和流程的能力。没有这种技能组合(无论是一个人还是多人),实现自动化就没有捷径。 令人兴奋的是,这在大多数企业中是一项全新的功能,将在经济中创造一大新角色。 每一代技术都创造了前所未有的就业机会,这些就业机会源于新兴行业的迫切需求。自动化工程师是这一代人从事这些工作的关键范例。目前还没有人有十年的工作经验,因为十年前制造这种可能的工具并不存在(甚至两年前也不存在)。 如果你具备软件技能并深入人工智能领域,这是一个需要深入探讨的领域。
Aaron Levie
Monster model right here
中文: 怪物模型就在这里
Aaron Levie
There's a huge opportunity right now in being the deployment layer for AI into the economy. The amount of work it takes to change out workflows in enterprises tends to be far greater than anyone realizes or would prefer. Clearly this is what the applied layer of AI is going to look like in the form of software and agents, but also it opens up new services firms opportunities. Legacy systems need to be moved to the cloud, data organization and access needs to be updated, software needs to be connected to agents in new ways, workflows need to be reengineered for agents, HITL needs to be figured out for the process, evals need to be generated and maintained, and the entire system needs to be continually updated as new models get released and new capabilities emerge. And the full list may even be longer. AI is not the same as just deploying software. Software you generally did the implementation of an existing, well understood category of technology, then stepped back and the customer kept running. With AI agents, you're delivering actual work augmentation to the organization, which has a completely different set of complexities associated with it. You're no longer deploying tools that the company is merely enabled by, you're deploying work output in a process. Completely different implementation and enablement process. As a result, this is going to open up lots of new kinds of firms and plays for existing firms to diffuse AI into organizations. We're going to see approaches by industry, by size of company, and by problem inside of companies. Traditional SIs will modernize and adapt (some will clearly not adapt as well), and new entrants will also be founded in this period that take advantage of this window. Great time to be an FDE or FDE firm.
Aaron Levie
It’s entirely plausible that the AI industry can deal with safety and security through a set of shared standards and practices for the foreseeable future. At some point in capability progress there will inevitably be greater oversight, testing, and layers of liability and regulation, but we are still in a phase where the industry can absolutely align on these efforts collectively. Awesome to see this come together so quickly and for the whole industry to support progress without reducing competition. This certainly points to a much brighter future for AI.
中文: 人工智能行业在可预见的未来通过一系列共享标准和实践来应对安全与保障,这完全是合理的。 在能力进步的某个阶段,将不可避免地加强监管、测试以及层层责任和监管,但我们仍处于一个阶段,行业完全能够共同推进这些努力。 很高兴看到这种情况如此迅速地走到一起,让整个行业在不减少竞争的情况下支持进步。这无疑为人工智能带来了更加光明的未来。
Aaron Levie
A good direction for the future of AI
中文: 人工智能未来的良好方向
Aaron Levie
At Box, we've been testing Sonnet 5.5 in early access on our complex work eval with the Box Agent, and Sonnet 5.5 is another strong jump from Sonnet 5 on knowledge work with enterprise content. Across the board we saw a 4 pt overall improvement on our hardest tests, but it was roughly 2.4× faster to a finished deliverable on 12% fewer tokens. We also saw major wins in key verticals, like a +18 percentage point improvement in Financial Services, +7 pt in Legal, +8 pt in Life Sciences, and +7 pt in the Public Sector. Here are some specific wins on certain tasks that show the power of Sonnet 5.5: * Financial services (+48 points): a due-diligence review of a trading-tool acquisition, where the deal book's arithmetic doesn't hold up. Sonnet 5.5 caught the miscalculated interest totals and the mispriced options and flagged them, and it also finished this work about 61% faster. * Legal (+14 points): a commercial lease review. Sonnet 5.5 refused to invent a "standard" market benchmark for renewal terms and notice periods, and instead flagged what the lease was genuinely missing. * Life sciences (+22 points): writing up a field trial from the raw measurement data. Sonnet 5.5 got the sample standard deviations right across both treatment groups. Sonnet 5 missed nearly all of them. And Sonnet 5.5 did it about 41% faster. * Public sector (+11 points): a report on a math intervention program. Sonnet 5.5 pulled the right student progress-monitoring figures out of the underlying data, about 71% faster. Customers will be able to build AI agents with Sonnet 5.5 in the Box AI Studio shortly.
Aaron Levie
Interesting to think about all the implications in a world where agents begin to make the best or most efficient choice for their users. On one hand, there is some subset of the economy that benefits from the friction customers traditionally have changing something about their habits. In those parts of the market, switching costs will come down and competition is going to increase dramatically until some equilibrium is reached between the disruptive alternatives and the incumbents. On the other hand, there are lots of markets that are hurt by a significant amount of friction, where agents will begin to unlock all kinds of economic activity by making it far easier to buy products or services that were too frictionful before. Healthcare, travel, local services, certain information services, and other entirely new markets probably are net beneficiaries of agents as a result. No matter what, a future meaningfully mediated by agents that work tirelessly for us and our goals can’t possibly function exactly the same as today. Going to be wild.
Aaron Levie
RT @a16z: Box CEO Aaron Levie, Steven Sinofsky, and Martin Casado join Erik Torenberg to discuss "We Must Pace the Frontier," the upcoming AI election in 2028, and Jev: They argue that most of today's AI regulation debate is happening before anyone has defined the risks being regulated. Every prior wave, from computer viruses to aviation, built its safety standards after learning how the technology actually failed. Then it gets concrete. Agents don't get tired, run at enormous scale, and probe systems in ways employees never could, which may mean rethinking permissions, authentication, and the security stack itself. They close on why AI innovation may increasingly happen outside the frontier labs, in the software built around the models. 00:50 "We Must Pace the Frontier" 03:50 Do the labs believe their own x-risk talk? 06:49 If it's existential, nationalize it 11:32 "You're asking us to regulate you?" 12:08 2028 as the AI election 18:50 Tech never learned to navigate regulation 25:50 The law that came from one 1983 hack 28:39 Noam Brown's heat exfiltration idea 32:30 Cold War covert channel stories 34:35 Agent swarms look like a DoS attack 41:25 Sinofsky's fear: GDPR for AI 46:20 No jets if the FAA started in 1910 48:38 Jev: decision engines vs. chatbots 53:01 Labs build beings, software needs tools YouTube: https://youtu.be/TLJNJDf2XGo @levie @stevesi @martin_casado @eriktorenberg
🎬
视频
Aaron Levie
You can’t automate what you can’t measure. This means that evals are one of the gates to diffusion of AI in the enterprise. We can test our deterministic processes through software, but most enterprises have no useful way of understanding how their non-deterministic processes are working today. Specifically the work that agents are doing for them. Evals are mission critical for enterprises adopting AI because you have no other way of knowing what’s working, what’s broken, what changed, what improved, what you can do more of, etc. if you don’t have a good sense of how agents work in your environment today. All changes, upgrades, and deployments are downstream from good evals. Not only are we going to get vastly more domain specific evals over time for the labs and across the industry, but every enterprise will also need a clear sense of how agents are performing in their environment as well. Huge opportunity.
中文: 你无法自动化你无法衡量的内容。这意味着,Evals 是人工智能在企业中传播的大门之一。 我们可以通过软件来测试我们的确定性流程,但大多数企业目前无法理解其非确定性流程的工作方式。特点是代理人员为他们所做的工作。 Evals 对于采用人工智能的企业至关重要,因为如果你对当前环境中的智能型工作方式没有充分了解,你无法了解哪些是工作、哪些是坏了、哪些是改进的,哪些是改进的,哪些是什么。所有变更、升级和部署均来自良好的evaval下游。 不仅随着时间的推移,我们还将为实验室和整个行业获得更多特定的领域,而且每个企业也需要清楚地了解代理在其环境中的表现。巨大的机遇。
Aaron Levie
@alexandr_wang @Box @Muse Imagine being a CIO and opening X right now
中文: @alerxand_wang @Box @Muse Imagine 成为首席信息官,立即开启X
Aaron Levie
Great vision for the future of the creative industry with AI from the one person who can actually predict these things. “As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world. In live action, filmmakers like Steven Spielberg, James Cameron and Peter Jackson embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process. Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process.” Technology has reshaped the creative industry over and over through the years, and has generally always led to an expansion of opportunity or new approaches as a result. And even when the techniques or mediums change, the underlying need for creative skills and taste don’t go away. Just more people are able to leverage those skills effectively when a new technology arrived, and users of the tools discover novel new ways to tell stories.
Aaron Levie
中文: 缪斯与博克斯终于在一起
Aaron Levie
What an insane day in AI. The frontier models just became substantially cheaper, with the Opus 5.5 price cuts, and now with GPT-6 Sol and Luna dropping token prices by 50%. The rate at which the cost per task (on a like-for-like basis) drops in AI is unlike any other type of technology in history. And every time the cost of AI drops, the use-cases you can deploy agents against dramatically increase. This is Jevons paradox applied to agents. These improvements will directly lead to broader diffusion of AI in the economy as we can use agents to process all of our data, scan our code for security issues, read through all log data to make decisions, have agent swarms in workflows, and much more. The cost of tokens is directly correlated to these use-cases being opened up at scale.
Aaron Levie
At Box, we've been testing Opus 5.5 on a variety of complex enterprise knowledge work tasks dealing with unstructured data with the Box Agent. Overall, we saw frontier capability levels, with major performance improvements over Opus 5. 63% fewer tokens used, 42% less verbosity, and 30% faster vs. Opus 5. And the model itself is cheaper, so this is a major win for any agentic computer use, coding, analytics, or data work that enterprises will be doing. Here are some examples of the task wins and performance gains across a variety of industry tests that we performed: • Financial services - due diligence (+39% task accuracy): A year of transaction records, with the job of finding every miscalculation in an acquisition target's pricing tool. Opus 5.5 scored a perfect result on every attempt in half the words Opus 5 used, consuming 82% fewer tokens overall. • Technology - cloud cost analysis (+65% task accuracy): Work out what a company should actually change about its cloud spend. Opus 5.5 picked the right basis for the retention calculation and kept the source data's unit conventions straight all the way through, so the number at the end actually holds up. It took half the time Opus 5 took, with 70% fewer tokens. • Consumer products - client account analysis (+17% task accuracy): Set the onboarding targets for a client account, reading across the signed contract, a satisfaction tracker and a team metrics sheet. The contract never states a senior/junior split, so Opus 5.5 derived it from the 18-person roster and showed the rule it used; several clients had a perfect 10 on individual survey questions, so it averaged each client's responses instead of crowning the single 10. It finished this one in half the time, on 78% fewer tokens. • Clinical diagnostics - data analysis (+15% task accuracy): Malaria rapid-test performance across a dry and a wet season: build the patient records out of two clinical PDFs, compute positive test rates by season and gender, and test whether parasite counts really differ between test-positive and test-negative patients. Opus 5.5 caught that the two groups' standard deviations differed more than 100-fold, re-ran it the right way, and found the dry-season difference didn't hold up after all. This accuracy gain came with a final answer that was half the length of Opus 5's, and also needed 78% fewer tokens end to end. Customers will be able to build AI Agents with Opus 5.5 shortly in the Box AI Studio.
Aaron Levie
The monetization potential of personal agents that are transacting on your behalf you is quite significant. If you imagine agents that are perfectly capable of handling an arbitrarily complex task end to end, then eventually a substantial amount of commerce inevitably goes through them. You’ll start by slinging your daily simple and annoying tasks at the agent. Then as people get used to it, they’ll just start to throw more complex tasks at the agent, ultimately leading to even more spend through these systems than what they were doing before. If you can bring down the friction for commerce and services, then you end up spending even more. This thus creates a ton of opportunity for the agent providers (Muse, etc.), but also completely new opportunities to build the layer that the agents want to interact with (commerce, local, b2b services, etc.). Win/win for multiple layers.
Aaron Levie
AI agents will use software 100X more than people ever did. Even as interfaces begin to fade into the background as you primarily interact with agents, those agents still need many of the core primitives that people have used. In fact, in many ways these core primitives become even more important when agents can take destructive actions in our systems or where the context they’re accessing with make or break the workflow. This will be true of where you house your CRM or ERP or structured or unstructured data platforms. The platforms that can best act as the security layer and guardrails for agents, manage the data for agents and people, and orchestrate the business logic for workflows have a huge opportunity right now. This is true for brand new startups as well as existing platforms that can move fast enough.
Aaron Levie
Literally impenetrable from agent swarms
中文: 从代理群中字面上无法穿透
Aaron Levie
Personal agents are the ultimate manifestation of “build something that agents want”. The form factor of a product like Muse is you want to be able to hand off a task to the agent and ensure that it is fully completed end to end. To do this, the agent must be able to successfully operate with your tools or use its own to complete the task. Use your MCP or CLI, easily navigate your site, be able to transact, and more. The new attention you need to compete for is not from the user itself but instead for the agent. This means that the tools that allow agents to order food, handle ecommerce transactions, book flights, work with the local economy, and interact with our data and information best, are the ones that will get used the most. This will ultimately be the biggest opportunity and shakeup in consumer tech since the App Store itself.
中文: 个人代理人是“建立代理人想要的东西”的终极体现。 像Muse这样的产品的外形尺寸是,您希望能够将任务交给代理,并确保任务端到端地完成。为此,代理必须能够使用您的工具成功操作,或自行完成任务。使用您的MCP或CLI,轻松浏览您的网站,能够进行交易,以及更多。 你需要争夺的新关注点,并非来自用户自身,而是来自代理。这意味着,让代理商能够订购食品、处理电子商务交易、预订航班、与当地经济合作,以及最好地与我们的数据和信息互动的工具,才是最需要使用的工具。 这最终将成为自App Store本身以来消费类科技领域最大的机遇和变革。
Aaron Levie
@alexandr_wang A literal human thought other literal humans were writing these posts? Incredible.
中文: @alexandr_wang 一种字面意义上的人类想法,其他字面意义上的人类正在撰写这些文章?难以置信。
Aaron Levie
Jev will be super helpful for agents to make split second decisions in workflows, data classification, judgment calls, and hundreds of other use-cases in the enterprise. Here's a quick demo with Box and Jev to make that real. The demo pulls an incident report from Box, asks whether it's customer-facing and how severe it is, moves the file into escalate, monitor, or review folders, and sets a metadata template instance with the result. This all happens nearly instantly and at almost no cost. You can imagine this in insurance claims, contract management, loan processing, security reviews, customer log analysis, and so on. Definitely a great new class of AI use-case.
🎬
视频
Aaron Levie
Agents already make up the majority of inference. This will quickly trend toward nearly all inference over the next year or two. The vast majority of tokens used in the world will be agents that are executing unbelievable amounts of tasks for us in the background 24/7. Agents will be deployed to read all code changes to secure our software, process all of our data inside of workflows, handle a significant majority of the research that goes into recruiting and customer prospecting, review every event stream and log from every system in the world, go out and execute tasks for us in our personal lives, and hundreds of other use cases. The rate of new agents coming online that will be consuming insane amounts of tokens is not slowing down. Just in the past week I’ve introduced multiple completely new workflows that would not have been possible technically even a month ago. Incredible time to be doing anything in inference and of course building on top of all of this.
中文: 代理已经占了大部分推断。这将在未来一两年内迅速趋向于几乎所有推断。 世界上使用的绝大多数代币都将是代理,它们全天候为我们在后台执行大量任务。 将部署代理来读取所有代码变更,以保护我们的软件,在工作流程中处理所有数据,处理大量关于招聘和客户勘探的研究,查看全球每个系统的活动流和日志记录,在个人生活中为我们完成任务,以及数百个其他使用案例。 新上线的代理数量将消耗大量代币,这种速度不会放缓。就在过去一周里,我推出了多个全新的工作流程,这些工作流程甚至一个月前在技术上都不可能实现。 是时候在推理上做任何事情了,当然也在这一切之上继续建设。
Aaron Levie
Incredibly exciting that there are entire universes of AI innovation that still exist that weren’t even on most of our radars. Being able to process information insanely quickly, at crazy low costs, with high levels of capability is huge for a wide number of enterprise tasks. Data classification tasks, routing decisions inside of a workflow, decision making when handed a particular domain problem, quick judgment calls about safety or security, and more all are the gates in a large number of processes. This model and approach could be quite cool in agentic workflows in the enterprise.
中文: 令人无比兴奋的是,仍有整个人工智能创新领域,甚至存在于我们大多数雷达上。 能够以极低的成本、高能力快速快速处理信息,对于众多企业任务而言,是巨大的。 数据分类任务、工作流程内部的路由决策、处理特定领域问题时的决策、关于安全或安保的快速判断,以及更多流程中的问题。 这种模式和方法在企业中的代理工作流程中可能相当酷。
Aaron Levie
RT @Box: Claude’s new Slides experience can turn deal materials in Box into an editable credit committee briefing. In our demo, @claudeai reviews a borrower’s deal workspace, creates a five-slide presentation with financial visuals and source references, and flags conflicting information for review. We leave a comment directly on a slide, Claude revises it, and the finished deck is exported as PowerPoint and uploaded back to Box through MCP for the credit team to review. Connect Box to Claude to create and refine presentations using your enterprise content.
🎬
视频
Aaron Levie
There’s a massive chasm between the power of AI models and the ultimate workflows that enterprises are trying to automate. This gap is the opportunity for the applied AI layer to fill. You need to connect the intelligence to workflows, often reengineer processes, aggregate the right context and data, allow for the right human in the loop experiences, drive change management, do domain specific evals, manage the security and governance of the data and process, and much more. We’re going to see this layer emerge in every vertical and horizontal category. And ironically, even as models improve at incredible rates, this layer still must exist - and may become even more important and useful. Greater capability enables even more complex tasks to be tackled, amplifying the challenges if you don’t do this well. Was super fun chatting with @sonyatweetybird on all the things going into AI diffusion.
中文: 人工智能模型的力量与企业试图自动化的终极工作流程之间存在着巨大的鸿沟。这一空白是应用AI层填补的机遇。 你需要将智能与工作流程连接起来,经常重新设计流程,汇总正确的语境和数据,在循环体验中实现合适的人,推动变革,进行特定的领域交流,管理数据和流程的安全与治理,以及更多。 我们将在每一个垂直和横向类别中都能看到这一层。具有讽刺意味的是,即使模型以惊人的速度改进,这一层仍然必须存在,并且可能变得更加重要和有用。增强的能力能够解决更复杂的任务,如果你做得不好,就会放大挑战。 和@sonyatweetybird聊天,聊到人工智能扩散的所有问题,这非常有趣。
Aaron Levie
RT @sonyatweetybird: Well.. @levie and I filmed this episode of Training Data a week or two ago, when the “current thing” was Doug Leone’s novacaine root canals instead of pacing the frontier… Simpler times! But Aaron’s advice on reinventing yourself and your company for AI is timeless. Aaron founded @Box 20 years ago. It sits on hundreds of billions of enterprise files, and he's bet the company on agents that can read every one of them. He's also one of the most wired-in people in AI, on every cap table and, by his own admission, 95% Twitter-educated. He’s the rare CEO who can straddle both the internet AND has the ear of CIOs. His core argument: (1) the gap between what a model can do and what an enterprise workflow actually needs is vast, and closing it is a lot of software; (2) diffusion of AI outside of coding will take far longer than Silicon Valley thinks, and that slowness is exactly where the applied layer's value comes from. The conversation covers: — why application companies are the hottest neolabs, and why the LLM-wrapper thesis is finally working — the fox-guarding-the-henhouse problem with letting model providers route your tokens — work slop, and why we accept AI-written code but flinch at AI-written decks — how Box built its agentic harness and why it beats raw API access on accuracy and latency — the open-weights paradox: closed labs and open models both growing exponentially at once — what continual learning has to solve before it works for a lawyer with five matters and a Chinese wall — why 90% of enterprise tokens in five years will come from tasks no human kicked off — the mandate for founders right now: whoever gets it to the customer wins 0:00 – Introduction 1:55 – Are application companies the hottest neolabs? 6:56 – Will the labs move up the stack? 12:34 – Box and betting the company on AI 16:50 – Hero use cases: reading a million contracts and long-running agents 18:42 – Work slop: why AI code is embraced but AI content isn't 24:08 – Building Box's agentic harness and the evals that matter 27:23 – The state of the model race 29:25 – Open-weight model adoption in the enterprise 32:34 – Memory, continual learning, and what belongs in the weights 37:29 – Box Labs and systems of record in a world of agents 44:55 – Will chat be the dominant UI for enterprise AI? 48:00 – Why coding diffused fast and the rest of knowledge work hasn't 54:31 – Staying wired in, making a company AI-first, and what it takes to win
🎬
视频
Aaron Levie
RT @matanSF: We have raised $200M at a $5B valuation to scale self-improving software development in the enterprise. @FactoryAI has grown to serve hundreds of thousands of developers at companies including RBC, Adobe, Nvidia, T-Mobile, and Palo Alto Networks. We will use this capital to accelerate our investments in research, product, and global go-to-market.
Aaron Levie
We probably need to all update our sense of what’s coming in terms of agentic workloads with the combination of agent swarms, better computer use, the next wave of APIs and MCPs coming online, and new form factors like Muse or Instinct, vertical enterprise agents, background workflow agents, and other products that are emerging right now. We’re going to throw agents at vastly more tasks in our professional and personal lives than we had initially imagined. The kind of information that agents will go out and find for us and do work for us in the background is going to be 100X more volume than what we can imagine previously prompting in a single session. You’re going to have agents go out and surgically recruit talent for you 24/7, look for every signal in your customer’s business for when to pitch them, process every single transcript and conversation for product insights, review every line of code written for security issues and bugs, brute force test your systems for issues, and 100s of other tasks. We’re probably 1% of the way into the shape of what all these agents look like, where they get deployed from, how they get managed, how we budget for them, and so on. But it’s inevitably going to happen at a scale that wouldn’t have been fathomable before.
Aaron Levie
Protecting enterprise data in a world where agents are using our systems 100X more than people ever did is going to be one of the more complex security and governance challenges of the 21st century. Importantly, security and productivity gains are inexorably linked in the world of AI. If you give an agent too much unfettered information access, it will be difficult to truly control and protect your data; and conversely, if you lock everything down completely, you won’t get any real productivity gains from AI. We need all new ways to protect our systems, environments, and structured and unstructured dada in the enterprise in an intelligent way by modernizing our approach to security and governance. At Box, as one example, we’re building new intelligent ways to protect enterprise data and agentic use of that information. A recent update in Box Shield is to provide granular controls on what content agents can and can’t work with based on document classification level. We’re also working on other features that can automatically detect and alert (or block) when data is being accessed or used in unusual or anomalous ways by agents. And this is just the start. There’s a ton more innovation coming across the entire industry - from the labs like OpenAI or Anthropic; security platforms like Palo Alto Networks, Cisco, CrowdStrike, Okta; or startups like Enclave, Method, Alterion, Runlayer, and many many others - to rethink how we protect information in the world of AI agents. Exciting and wild times ahead.
Aaron Levie
“Pacing” can be somewhat of a trigger word because it sounds like an arbitrary slow down of capability or a way to hobble competitors through undue regulation. However, the specific improvement goals laid out in Dario’s are absolute necessities in AI development. We expect equivalent in areas like aerospace, life sciences, health care, and other industries, and AI development likely shouldn’t be that different. AI is going to be the technology underpinning our financial trading systems, our medical devices, our biotech breakthroughs, defense systems, government workflows, and many more mission critical domains. So it completely stands to reason that we want these systems to be safe and “aligned”. How we get there -without meaningfully slowing down innovation progress or reducing competition- is one of the most complex questions in the 21st century, but the need is clearly real.
中文: “节奏”可能在某种程度上是一个触发词,因为它听起来像是能力的随意放缓,或是通过过度监管来阻碍竞争对手的一种方式。 然而,达里奥提出的具体改进目标是人工智能发展中的绝对必要条件。我们预计在航空航天、生命科学、医疗保健和其他行业等领域将出现相应的类似状况,而人工智能的发展可能不应如此不同。 人工智能将成为支撑我们金融交易系统、医疗器械、生物技术突破、国防系统、政府工作流程以及更多关键任务领域的技术。因此,我们完全有理由希望这些系统安全且“一致”。 我们如何实现这一发展——在不显著减缓创新进步或减少竞争的情况下——是21世纪最复杂的问题之一,但需求显然是真实存在的。
Aaron Levie
“Pacing” can be somewhat of a trigger word because it sounds like an arbitrary slow down of capability or a way to hobble competitors through undue regulation. However, the specific improvement goals laid out in Dario’s are absolute necessities in AI development. We expect equivalent in areas like aerospace, life sciences, health care, and other industries, and AI development likely shouldn’t be that different. AI is going to be the technology underpinning our financial trading systems, our medical devices, our biotech breakthroughs, defense systems, government workflows, and many more mission critical domains. So it completely stands to reason that we want these systems to be safe and “aligned”. How we get there -without meaningfully slowing down innovation progress or reducing competition- is one of the most complex questions in the 21st century, but the need is clearly real.
中文: “节奏”可能在某种程度上是一个触发词,因为它听起来像是能力的随意放缓,或是通过过度监管来阻碍竞争对手的一种方式。 然而,达里奥提出的具体改进目标是人工智能发展中的绝对必要条件。我们预计在航空航天、生命科学、医疗保健和其他行业等领域将出现相应的类似状况,而人工智能的发展可能不应如此不同。 人工智能将成为支撑我们金融交易系统、医疗器械、生物技术突破、国防系统、政府工作流程以及更多关键任务领域的技术。 因此,我们完全有理由希望这些系统安全且“一致”。我们如何实现这一发展——在不显著减缓创新进步或减少竞争的情况下——是21世纪最复杂的问题之一,但需求显然是真实存在的。
Aaron Levie
Good post. Don’t agree with all of it, but it represents many of the actual practical realities for the general path forward in frontier AI whether we like it or not. At the level of capability we’re seeing from models there will inevitably be some form of coordinated self-regulation of the industry; that’s generally a good thing. The big question will be if everyone can agree to what that is. However, the way things are headed politically there’s a chance that the labs won’t even get a say in it. And then the bigger question yet is do all the countries participate or not. Any “slow down” hinges on broad participation, which from a game theory standpoint isn’t likely until the risks are more severe and obvious. Expect this to be pretty messy for a while.
中文: 好文章。并非所有情况都同意,但它代表了我们是否喜欢前沿人工智能领域未来总体路径的许多实际现实。 在我们从模型中看到的能力水平,行业中不可避免地会出现某种形式的协调自我调节;这通常是一件好事。最大的问题是,是否每个人都能同意那是什么。然而,政治上的发展方向,实验室甚至可能对此没有发言权。 那么更大的问题是,所有国家是否都参与其中。任何“慢下来”都取决于广泛参与,从博弈论的角度来看,这种参与在风险更加严峻和明显之前不太可能。 预计这种情况会相当混乱一段时间。
Aaron Levie
Good post. Don’t agree with all of it, but it represents many of the actual practical realities for the general path forward in frontier AI whether we like it or not. At the level of capability we’re seeing from models there will inevitably be some form of coordinated self-regulation of the industry; that’s generally a good thing. The big question will be if everyone can agree to what that is. However, the way things are headed politically there’s a chance that the labs won’t even get a say in it. And then the bigger question yet is do all the countries participate or not. Any “slow down” hinges on broad participation, which from a game theory standpoint isn’t likely until the risks are more severe and obvious. Expect this to be pretty messy for a while.
中文: 好文章。并非所有情况都同意,但它代表了我们是否喜欢前沿人工智能领域未来总体路径的许多实际现实。 在我们从模型中看到的能力水平,行业中不可避免地会出现某种形式的协调自我调节;这通常是一件好事。最大的问题是,是否每个人都能同意那是什么。然而,政治上的发展方向,实验室甚至可能对此没有发言权。 那么更大的问题是,所有国家是否都参与其中。任何“慢下来”都取决于广泛参与,从博弈论的角度来看,这种参与在风险更加严峻和明显之前不太可能。 预计这种情况会相当混乱一段时间。
Aaron Levie
Great time to pivot to software
中文: 转向软件的绝佳时机
Aaron Levie
This one is very cool. Now you can mount Box to agent sandboxes to make it far easier for an agent to read and write files on the agent's computer. As AI agents execute critical workflows in the enterprise, they're going to need the same primitives that people have had.
中文: 这个非常酷。现在您可以将Box安装到代理沙盒,以便代理在代理计算机上读取和写入文件变得更加容易。当人工智能代理在企业中执行关键工作流程时,他们将需要与人们相同的原语。
Aaron Levie
Some more tales from the road. Met with a couple dozen technology leaders this week across banking, media, information services, insurance, and consulting to discuss agents in the enterprise. Some of the biggest trends right now: * Cyber! Everyone nervous about the growing rate of vulnerabilities coming at them from AI, and the implications of the OpenAI Hugging Face incident. The conversation is not as existential as it is in Silicon Valley, but still highly concerned and pragmatic about what to do about it operationally in their environments. Lots of new discoveries due to AI, and still hard to keep up with all the changes they have to execute now. * Model battles persist. Most companies are deploying multiple frontier models within their enterprise. Too hard to standardize on anything and seeing different preferences across their teams and use cases. But the dollars are still concentrated on just a few vendors. Open weights still in infancy at scale in most of these organizations, often due to lack of domestic “frontier” OSS options. Plenty of appetite for more options here, but so far few places to go. * Agent security and identity. Somewhat tied to Hugging Face, there’s much more awareness to the new challenges around agent security and identity management in a world when agents are trying to get into every system they can. In a perfect world enterprises could setup identities for all their agents and control what they’re doing, but of course sometimes the agent needs to act exactly as the user as well. * Process reengineering. Most companies realizing that the big upside of agents is when they can change the actual workflow itself to get the full gains from AI. Far more ROI when companies can adjust their workflows to support agents changing how the work happens instead of just layering on agents into the existing flow. But the big question is who can actually tackle driving these changes, where does that live, etc. Best lessons were still around embedded FDEs in the functions. * Ruthless adjusting of architectures. Most companies had examples of changing systems out multiple times just in the past year or two with different vendors. I probably haven’t heard “we tried X and it didn’t work so have gone with Y” more than in today’s environment. The lesson here is that because innovation is happening so fast, no one hangs around until a vendor gets something right, they just move on to the next one. * Evals! Still very early for most companies to have a good grasp of evals of their workflows. A few customers out of a couple dozen called this out - huge opportunity right now for enterprises to have a good sense of how their work actually happens and how well AI is doing against it. * Legacy systems still a hurdle. As always, legacy systems still remain a mainstay issue that holds back enterprises from rapid adoption of AI in enterprises. Data is fragmented across legacy environments that weren’t built for an agentic world. Companies spending a lot of time just cleaning up these old platforms. Many more topics, but these tend to be some of the more top of mind items at the moment in the enterprise.
Aaron Levie
Excited to partner more deeply with OpenAI so you can work with your enterprise content from Box securely in ChatGPT. Software continues to go headless, and we’re massive believers that the future of work will be AI agents that can process data and execute workflows anywhere.
中文: 很高兴能与OpenAI更深入地合作,以便您能够安全地在ChatGPT中使用Box的企业内容。 软件仍在持续无奈,我们坚信未来的工作将是能够在任何地方处理数据并执行工作流程的人工智能代理。
Aaron Levie
The world is going to be using coding agents for far more than anyone would have thought before. We’ll be building tools in all new categories, building software for companies that couldn’t afford it before, using agents to protect systems from cyber risk, automating research in life sciences, processing insane amounts of data in complex workflows, upgrading legacy systems and infrastructures, and hundreds of other use cases. In short, by lowering the cost of code, we are going to do so much more with it. This will result in software becoming more useful and the leverage of engineers going up dramatically, which coincidentally will mean we need more of them, not fewer.
中文: 世界上使用编码代理的使用将远远超过之前任何人的想象。 我们将在所有新类别中构建工具,为以前负担不起的公司开发软件,使用代理保护系统免受网络风险的影响,自动化生命科学领域的研究,在复杂工作流程中处理大量大量数据,升级传统系统和基础设施,以及数百个其他使用案例。 简而言之,通过降低代码成本,我们将用它做更多的事情。这将使软件变得越来越有用,并极大地利用工程师的杠杆作用,而巧合的是,我们需要的是更多,而不是更少。
Aaron Levie
The way to the reconcile capability level of AI vs. GDP impact is that the diffusion of AI will take much longer than people think. And it will also show up in ways that are hard to measure in GDP immediately. You could bring the world’s greatest superintelligence to many workflows, and still be bound by the laws of corporate physics: getting data prepared and put into a pipeline, process reengineering and change management, aligning on how the new workflow should function, and so on. Even after you solve all that, you’re still bound by the speed of the real world: waiting for a customer to respond to a proposal, getting a permit for a project, a drug discovery pipeline taking years to eventually reach the consumer, and so on. Not to mention lots of positive daily AI use-cases are entirely net neutral to GDP, at least in the near term. AI diffusion is going to be the theme of the next decade. The upside is that there’s a tremendous amount of opportunity in building the bridges between superintelligence and real-world workflows.
中文: 实现人工智能与能力水平对调的方法。GDP的影响在于,人工智能的传播时间将比人们想象的要长得多。而且它将以难以立即衡量GDP的方式出现。 你可以将全球最伟大的超级智能引入多种工作流程,并且仍然受企业物理定律的约束:将数据准备并投入流程,进行流程再造和变革管理,对新工作流程的运行方式进行对齐,等等。 即使解决了所有问题,你仍然被现实世界的速度所束缚:等待客户回复提案,获得项目许可,以及药物发现管道,最终需要数年时间才能联系到消费者,等等。更不用说许多积极的每日人工智能使用案例,至少在短期内完全对GDP保持净中性。 人工智能的扩散将成为未来十年的主题。好处是,在构建超级智能与现实世界工作流程之间的桥梁方面,存在巨大的机遇。
Aaron Levie
@alexandr_wang 😮
Aaron Levie
Personal assistant agents are going to be a very exciting AI category. It’s the first time you can have high token volume agentic use-cases that make sense for consumers. Lots of different approaches emerging right now, and it’s going to be hyper competitive because these agents will mediate a lot of consumer spend over time. But this certainly plays directly to Meta’s strengths: lots of compute required, can monetize with ads and commerce, software focused experiences so can distribute it at scale, and so on.
中文: 个人助理代理将成为一个非常令人兴奋的人工智能类别。这是您首次能够拥有对消费者而言具有高代币交易量的代理用例。 目前出现了许多不同的方法,而且竞争将非常激烈,因为这些代理会随着时间的推移大量消费。 但这显然直接发挥了Meta的优势:需要大量计算,可以通过广告和商业实现盈利,以软件为中心进行体验,以便大规模进行分发,等等。
Aaron Levie
Good reminder here that with rapidly accelerating AI progress, you should likely be building with a vision that contemplates at least a few orders of magnitude of capability improvement or token volume available to you. The best opportunities right now are those where you can value to customers today with things that are barely only possible, but with a full mission that feels nearly impossible with today’s tech. There are lots of products the require orders of magnitude more processing of data (and thus are barely economically viable ideas) or where model capabilities still are in their infancy relative to what’s needed to work. Best to be building for those scenarios.
中文: 提醒您,随着人工智能加速快速发展,您可能需要具备至少几级功能提升或代币容量的构想。 目前最好的机会是,如今你只需用几乎不可能的东西,就能为客户带来价值,而这些东西却充满了完全的使命,而如今的技术几乎无法实现。 有许多产品需要对数据进行更多数量级的处理(因此几乎无法实现经济可行性),或模型能力相对于需要的需求仍处于初级阶段。最适合为这些场景进行构建。
Aaron Levie
The AI jobs prediction so far is playing out the opposite of what many thought. The reason this is happening is because AI is bringing automation to areas that don’t have finite demand characteristics. AI allows companies to do far more than they could before, but agents still require human operators and oversight to generate value from them. And not only has AI not replaced many of the types of jobs that theoretically would have been most affected by AI - engineers or lawyers - but AI is also creating or growing entire categories of jobs across the economy. Cybersecurity, FDEs, agent operators, engineers in non-software domains, and tons of job categories are going to explode. It’s not clear if or when this trend will reverse because we keep discovering all new areas of demand that AI is going to open up.
中文: 到目前为止,人工智能的就业预测与许多人的想法恰恰相反。 之所以发生这种情况,是因为人工智能正在将自动化引入那些没有有限需求特征的领域。人工智能使企业能够比以往做得更多,但代理机构仍需要求人工操作员和监管人员从中创造价值。 人工智能不仅没有取代许多理论上受人工智能影响最大的工作类型——工程师或律师——而且人工智能也在经济领域创造或扩大了整个就业类别。 网络安全、FDE、代理运营商、非软件领域的工程师以及大量工作类别将会激增。目前尚不清楚这一趋势是否会或何时会逆转,因为我们不断发现人工智能将开放的所有新需求领域。
Aaron Levie
The internet is almost entirely unprepared for a future where everyone’s personal agents are running around executing tasks for them. There are going to be endless new challenges and opportunities at the infrastructure layer, for user experiences, new business models, and more.
中文: 互联网几乎完全没有准备好迎接未来,每个人的个人代理人都在为他们执行任务。 基础设施层面将出现无限的新挑战和机遇,这些挑战和机遇将带来用户体验、新的商业模式等。
Aaron Levie
If agents produce the vast majority of software in the future, and they’re most trained on open source software, they will inevitably do their best work with those tools. If you cycle this enough times, it means that open source effectively becomes the dominant software in the future as it’s what everything important gets built with. This was already the trend in many critical domains, but agents will accelerate this far faster than humans ever could have.
中文: 如果代理未来能生产出绝大多数软件,并且在开源软件上受过最大培训,他们必然会在这些工具上尽自己最大的努力。 如果你循环使用这些时间,意味着开源在未来实际上成为主导软件,因为它正是所有重要事物所构建的。 这在许多关键领域已经是趋势,但代理机构将比人类所能实现的更快加速。
Aaron Levie
Good post on how to think about process redesign in an enterprise with AI Agents. Unfortunately for most workflows, there’s no “easy” button. The best way to get real efficiency gains is you have to actually reengineer the process to take advantage of what agents are actually good at and can do differently and faster than people. One of the challenges of course is that the most important workflows actually space many functions at once, so someone needs to be able to go in and retool how work happens. “Nobody in that chain is empowered to walk into finance, for example, and say that their 14 step process should actually only be 5 steps. So nobody does it. Instead, you 'apply AI' on what you currently have, and you end up with faster sh*t.” This is the road ahead for AI diffusion in enterprises. This also presents the clear opportunity for companies - both at the applied AI layer (regardless of the article’s title) and teams or companies that can go and implement transformation into the enterprises. And all roads lead to FDEs one way or another.
Aaron Levie
@OpenAIDevs Great Model!
中文: @OpenAIDevs 很棒的模特!
Aaron Levie
Palo Alto Networks for message boards is going to be a $100B company
中文: Palo Alto Networks 的留言板将是一家价值 1000 亿美元的公司
Aaron Levie
GPT-6 Astra is out. We've been testing the model in early preview on our enterprise complex work eval at Box. It is now the best model we've ever tested on our expanded and hardest test set. Overall, GPT-6 Astra offers a breakthrough level of capability in coding, analytics, logic, and domain specific knowledge for dealing with complex enterprise knowledge work. In our eval, we use the GPT-6 Astra model in the Box Agent and give it a set of difficult tasks that represent real-world work in various industries. These tasks often will take people hours to execute properly and require deep domain expertise. GPT-6 Astra scored 77% overall vs. 74% with GPT-5.6 Sol, representing frontier performance. But the story is in some of the bigger gains on individual task types that are incredibly complex across. Here are a few examples across the tests that show how much of an impact this will be in different industries: * Media and entertainment (48% → 100%, +52 points). Ranking film genres and countries by profitability across a year of production data from four teams, where the trick is to apply the following year's tax-incentive corrections without double-counting them. GPT-6 Astra was perfect on every attempt. Sol had the rankings right but the underlying ratios were wrong. * Technology (69% → 97%, +28 points). Choosing which region to fund first, from a stack of performance and infrastructure documents that never state the metric the brief asks for. GPT-6 Astra flagged the gap, labelled its own figure a proxy, and caught a growth claim that didn't match the numbers beneath it. GPT-5.6 Sol reported the proxy as the real thing. * Legal (69% → 93%, +24 points). Reviewing an NDA against a company's own contracting policy, where reaching the right verdict isn't enough; the answer has to cite the provision behind it. Both models declined to approve the draft, but only GPT-6 Astra separated whether the liability cap's structure was permissible from whether its amount was defensible, and pointed to the policy language that settles it. GPT-5.6 Sol argued the same conclusion without citing the provision, which is a critical error in the legal industry. * Healthcare (53% → 77%, +23 points). Auditing a batch of radiology reports for errors and rating how serious each one is. GPT-6 Astra caught a terminology error in a knee-imaging report that GPT-5.6 Sol missed on both attempts, and was more reliable at grading severity rather than just flagging that something was wrong. * Energy (82% → 97%, +15 points). Building a consumption report on two facilities from a year of meter logs, where only one of the two sites is actually missing data. GPT-6 Astra found it and drew a line most reviewers wouldn't: the other site's odd solar readings are a measurement problem to investigate, not a gap in the record. GPT-5.6 Sol reported both sites as incomplete. Astra clearly is going to offer a meaningful jump in powering and orchestrating enterprise workflows. We'll be making GPT-6 Astra as an option in the Box AI Studio for customers to build agents with shortly as it continues to roll out.
Aaron Levie
Another huge moment for open weights AI. The infra platforms are working. The models are getting better. The ecosystems are being deeply invested in. The business models are working.
中文: 开放权重人工智能的又一重要时刻。基础设施平台正在发挥作用。模型正在变得越来越好。生态系统正在被深度投入。商业模式正在发挥作用。
Aaron Levie
This week is absolutely nuts for AI releases, and now a major new model update from Muse. We appear to be reaching escape velocity in capability progress. If this thing gets released as open weights it completely changes the dynamic of US open weights competitiveness.
中文: 本周对人工智能发布来说绝对是疯狂的,如今缪斯推出了一项重大的新模型更新。我们似乎在能力进展方面达到了逃逸速度。 如果将此功能以开放权重释放,将彻底改变美国开放权重竞争力的格局。
Aaron Levie
RT @danlovesproofs: Eight months ago we launched a bug finder that you could run once a month. Today we're launching Self Healing for your codebase. Introducing Detail! https://twitter.com/danlovesproofs/status/2095182189499711759/video/1
中文: RT @danlovesproofs:八个月前,我们推出了一个可以每月运行一次的漏洞查找器。今天,我们将为您的代码库推出自我修复。 介绍细节!
🎬
视频
Aaron Levie
AI for cyber is about to go vertical. The models increasingly becoming insanely good at finding and exploiting vulnerabilities. Frontier models are ahead, but we’re already seeing that open weights is not far behind. Most enterprises are already inundated with cyber discoveries, so now that’s only going to multiply. Triaging and automating the fixes with more AI -along with human oversight- is essentially the only way forward. In case you were wondering what jobs AI was going to create, it’s definitely going to great time to be in security.
Aaron Levie
At Box, we’ve been testing Fable 5.1 in early release against our complex enterprise work eval. Fable 5.1 delivers a huge 7 percentage point jump over Fable 5 for unstructured data tasks in the enterprise. On this updated test, we use the Box Agent with Fable 5.1 to work through a wide range of real world enterprise scenarios with documents in financial services, life sciences, the public sector, and more industries. All of these tasks require a high degree of analytical, math, logic, and domain knowledge to perform successfully. Here are a few examples of wins: * Financial Services (+17% improvement): on a tax-adjusted profit projection, Fable 5.1 correctly applies capital allowances before computing tax liability; this is a subtle ordering that Fable 5 misses, producing wrong figures all the way through retained earnings. * Technology (+37% improvement): on a cost-optimization analysis where a key metric's normalization is ambiguous, Fable 5.1 recognized the ambiguity, computed both forms, and presented the correct one, where Fable 5 committed to the wrong normalization. * Public Sector (+16% improvement): on an educational data analysis task, Fable 5.1 works through the full weighted-mean methodology and produces correct rankings, while Fable 5 miscategorizes one item and cascades errors through the whole sheet. These are just a few of the wins we saw. Overall major jump in capability for long running agentic workflows in the enterprise. Fable 5.1 will be available shortly in the Box AI Studio for building custom AI agents with enterprise content.
Aaron Levie
RT @drfeifei: I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀 Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
中文: RT @drfeifei:我非常兴奋,我们的@theworldlabs团队今天实现了一个重要的里程碑!推出Atlas——首个从零开始训练的多模式世界模式!🚀 Atlas 能够利用像素完美的相机控制生成帧,从单个输入图像中重建大场景,通过重新帧化视频模拟时空,从一个或多个输入图像中本地输出3D空间,将多个摆正图像组合成一致的3D世界,以及更多!这是有史以来最好的相机条件型世界型号,为从视觉特效(VFX)到机器人技术的多种可能使用案例打开了大门。我为我们的团队感到非常自豪!♥️
Aaron Levie
Now that the base open weights AI models are getting far better, and post training infra is becoming more mature and commercialized, there are going to be all new plays for companies that have large amounts of data to have their own models. Licensing data for external model training was previously the only play if you had a large corpus of information, but you can now reasonably go and train your own models as well without incurring the cost and complexity of competing with the labs on research. The general purpose frontier models will still have a leg up in broad areas due to the ability to handle the widest set of tasks, but you can absolutely see a future where we have far more models than today across every vertical and domain.
中文: 如今,基础开放权重的人工智能模型正在变得越来越好,而后训练的红外模式也正变得越来越成熟和商业化,对于拥有大量数据、拥有自身模型的公司而言,将会出现所有新的游戏。 使用外部模型训练的许可数据,以前是您拥有大量信息库的唯一选择,但现在您可以合理地自行训练模型,而不会产生与实验室在研究上竞争的成本和复杂性。 通用前沿模式在广泛领域仍具有优势,因为能够处理最广泛的任务,但你完全可以看到未来,我们在每一个垂直领域和领域拥有比现在更多的模型。
Aaron Levie
As AI security events pick up, it’s going to be critical that we have the most sophisticated AI agents for being able to detect and prevent security issues. Frontier models are clearly still ahead in cyber, but open models are catching up quickly. Wild times.
Aaron Levie
Sometimes people don’t have an intuitive sense of what Jevons paradox looks like for token consumption. Those of us working with enterprises get to see this first hand every day. Basically enterprises have an unending stream of tasks that they’d love to be able to bring automation to. But for each individual task, it’s either ROI positive or not based on the cost of bringing automation to it (including the setup time, change management, and so on). As tokens get cheaper (at a certain capability threshold), enterprises can afford bringing them to more of the work that they do. This could be processing every contract, reading every log, watching new streams of data for insights, having background agents execute workflows, and so on. Any time we can lower the cost of tokens, we will see a disproportionate increase in consumption. Even a 50% drop in token prices could result in a 5X increase in tokens for these kinds of workloads. This is why it’s critical to keep bringing down the cost of AI, and why that’s good for all market participants.
Aaron Levie
RT @Teknium: Take advantage of @Box in Hermes Agent to jumpstart your business!
Aaron Levie
The average strongly held belief these days in AI has a half life of 6 months at best. Here are just a few the industry has cycled through and probably has no consensus on at the moment: * OSS is too far behind to catch up * The labs can’t be profitable at scale * All software will be replaced by agents * You can’t build moats on top of models * Cheaper models will mean less compute * You don’t need evals * RAG is dead * AI will decimate engineering jobs * Prompting won’t matter in the future * We’ve hit a training wall * Frontier models are too dangerous to release …and dozens and dozens more. The key here is to remain flexible in your thinking because we’re going to be in a constant state of change for a while.
中文: 如今,普通人坚信人工智能最多半个月的寿命是半月。以下是行业周期中仅有几个,目前可能尚未达成共识: OSS 落后太远,无法跟上 实验室无法大规模盈利 * 所有软件将由代理人员更换 * 无法在模型上构建护城河 更便宜的模型意味着更少的计算 你不需要薎 拉格死了 人工智能将摧毁工程岗位 提示在未来并不重要 * 我们遇到了训练障碍 前沿模型太危险,无法发布 还有几十个,还有几十个。关键是保持你的思维灵活性,因为我们将暂时处于持续变化的状态。
Aaron Levie
This week’s tech earnings calls were a critical reminder of how valuable the relationship is between software and AI. Software provides the guardrails for how data is managed, how business logic in workflows are maintained, how you manage and govern access to information, how you protect data, and more. You want these systems to work the same way every time, without fail. Agents, on the other hand, work inside these systems and with this data to execute on tasks just like people would have. But at a scale far greater than people ever did, which is especially why the deterministic controls matter more than ever. And in fact many of the best ways to deploy agents will be directly within the software systems themselves, like from Salesforce, Box, Harvey, ServiceNow, and so on. These platforms can optimize agents for their workflows and domains, with evals and context that are unique to their vertical, as well as connect to other agents headlessly. The net result is both software and AI adoption go up, hand in hand. Software and Agents are going to work together, and ultimately grow the size of the IT TAM dramatically as a result over time.
Aaron Levie
Yesterday, at Box we reported our Q2 results with revenue of $321.1 million for the quarter, up 9%, or 11% in constant currency (our highest constant-currency growth rate in 14 quarters), and we raised our full-year revenue target to $1.290 billion. Overall, this growth is being driven by the demand we're seeing from enterprises aiming to get the most out of their enterprise content and transform in the era of AI. As I shared on our earnings call, during the quarter I spoke with many enterprise technology leaders who highlighted their primary goals and challenges in implementing AI. One of the most common topics is how enterprises can get the right context to AI agents, as well as tap into the full value of their unstructured data, in a secure and governed way. The world’s most advanced superintelligence is only as useful as the underlying enterprise knowledge it has access to. Agents need access to the critical information that makes up the bulk of corporate knowledge --like contracts, research materials, financial documents, marketing assets, product roadmap information, and so on-- most of which lives within an enterprise’s unstructured data. We've long been able to query, calculate, and automate work that deals with structured data; but the vast majority of data in an enterprise is unstructured, often in the form of our enterprise content. Instead of sitting on millions or hundreds of millions of files they know very little about, companies can now use AI agents to ask questions about this data, mine it all for intelligence, and automate nearly any workflow that involves this enterprise content. The technology leaders I'm talking to are also recognizing that as AI model capabilities advance across AI labs --like OpenAI, Anthropic, SpaceXAI, Meta, Google, NVIDIA, and a variety of open-weight providers-- they will need applied AI platforms that can bring the full power of these models to their processes. And with AI token usage continuing to grow exponentially, the ability to draw the right cost-performance mix becomes essential. Rather than migrating data and workflows into separate systems to unlock AI's benefits, enterprises need the ability to swap models or agents on their workflows and information at any time. Finally, as we saw with the recent OpenAI Hugging Face incident, enterprises will increasingly need platforms that can securely protect their corporate data and ensure that neither humans nor agents can get access to information they shouldn’t have access to. As external AI agents interact with enterprise data, enterprises will need platforms that offer robust guardrails, comprehensive audit logs, and real-time security alerts to ensure content remains protected and governed at all times. Given all of these dynamics, there's going to be a tremendous amount of opportunity for various plays to bring intelligence to the critical workflows inside of enterprises at the applied AI layer. At Box. we're focused on building the leading intelligent content management platform to empower enterprises to transform how they work with their unstructured data at scale. And we're looking forward to partnering with all of the other companies bringing intelligence into organizations as well.
Aaron Levie
Good post on what the applied AI strategy looks like at scale. It’s clear that there’s a wide gap between the AI models and the underlying workflows of an enterprise, which leaves a ton of opportunity for applied AI companies. “The world doesn’t just want raw models and agents; it wants problems resolved and outcomes achieved. The premium will sit with the companies that can diffuse this intelligence through every aspect of civilization, converting raw tokens into real world outcomes, transforming industries, and creating economies in the process.” This requires understanding the context, driving the change management, having a harness that can route to various models, connecting to the critical business systems in that vertical, solving the UX challenges of connecting users to agents in the right way in a workflow, understanding the evals in the space, and so on. That’s a ton of value that goes beyond just the model intelligence itself. And there’s a window of opportunity right now to build the defining companies that can bring intelligence to the critical domains in an enterprise.
Aaron Levie
Systems of record have never been more important than in a world where you have AI agents that will do 100X more work on these platforms than people ever did. Agents will be querying the data in these systems, processing tasks, executing workflows, collaborating with human and agent users, and more. Thus, the governance, reliability, security, access controls, and business logic remain critical more than ever — the OpenAI Hugging Face incident is just a small peek into what the future will look like with agents running around trying to execute on their goals. But this only works for the systems of record that respond with the right product experiences, APIs, and business models to support agents on and off their platform executing in these systems. Huge moment of opportunity and disruption right now in all of software.
Aaron Levie
It’s not fully appreciated that ZDR is responsible for a substantial amount of the growth of AI, because it dramatically simplified the compliance process most companies deal with in handling data with subprocessors. The alternative path would have been a far slower adoption rate. For the applied AI tools that are primarily responsible for driving AI growth, they generally have agreements in place with their customers to only offer models with ZDR to ensure that customers don’t have to inspect each model they turn on. Anything else is an exception process that takes far longer. And then inside enterprises directly, most companies have internal governance requirements that say they can only use ZDR models because they have no way of explicitly separating PII and other sensitive or confidential data they put in the context window across their enterprise. Now, all of this could change if the entire industry went this direction, but with one off models hold outs it would take years for these policies to change. Without ZDR, AI diffusion grinds to a halt.
Aaron Levie
Aaron Levie
AI diffusion is far more rate limited by having good evals than most realize. The kind of evals that you see for every model release are incredibly helpful, but only tell you the shape of general AI progress and the relative capability level of models. The far bigger space over time are evals on all the major workflows that enterprises do, down to the specifics of an individual company. This will be a huge space over time because you can’t automate what you can’t assess the progress on. Enterprises will not be able to go just on vibes.
Aaron Levie
The rate of progress in AI right now is unlike any other period in tech history. Models are getting cheaper on a like for like task basis, more generally capable, faster, and they’re going deeper in almost every domain. As intelligence becomes too cheap to meter, then the big opportunity is driving diffusion of AI into the economy. This is a great time to be applied AI company because there’s an incredible amount of innovation and competition acting as a tailwind to your own progress. Huge moment for startups right now that can take advantage of this.
Aaron Levie
Great post on what post training looks like for applied AI use-cases to bring down costs and improve accuracy on certain tasks. This will increasingly be an approach that companies that can get closer to the underlying workflow in an enterprise will take. The key is that once you understand a domain well enough and have enough volume on a set of similar tasks, it can start to make sense to purpose design models just for that work. “In post-training, we incentivized efficient tool use and reasoning through reward shaping, preferring trajectories that would reduce tokens consumed at inference-time given equivalent performance. This allowed us to co-optimize for both cost and quality, gaining significant performance while keeping cost stable.” Now, this won’t make sense in every domain, as general purpose frontier closed or open models will be good enough out of the box -or necessary- for the work at hand. But once you have deep enough vertical expertise, and either the costs are too high to do at scale *or* you have a unique enough task type not being trained on otherwise, this will make a ton of sense. Very compelling value proposition for being an applied AI company, and awesome to see multiple paths to winning in the market right now.
Aaron Levie
One of the more interesting dynamics with AI agents is trying to find the right form factor for how they should show up inside a workflow. As agents can do longer and longer tasks and need less intervention, they can show up more in the tools that we already use when working with other people or processes. This wasn’t possible when you had to constantly oversee the work they were doing, but becomes much more plausible when you send off a task and just see the work that is done at the end (or in chunks). Of course this will mean most software has to evolve to support this style of interaction, introducing new components and patterns to make this agent interaction work well. So it’s very cool to see what Slack is doing on this front. We can expect to see much more of this across platforms, and it will define which ones end up making sense in the future and those that don’t.
Aaron Levie
There tends to be a debate between being an expert or generalist in the era of AI. So far, the experts appear to have the upper hand, and that’s not slowing down. AI makes it 10X easier to get started with any kind of task. Coding, legal work, research, financial analysis, and so on. But having the right judgment for how to direct the agent onto the right work, how to veer it to course correct it, being able to review or test the output, and having the right sense of what “good” looks like all requires a high degree of skill depending on the field. That skill can be developed far faster due to AI for anyone interested, which is amazing for uplifting anyone in their career. But there’s no replacement to needing judgment and many of the core skills that go into most domains of work. And if anything, AI will be a technology that exacerbates differences in skill levels because the experts have far more leverage than ever before. Net net: don’t give up on being an expert at something.
Aaron Levie
Good details on the Stripe + OpenRouter deal here. For AI to diffuse more broadly, developers and enterprises will want ways of being able to mix and match intelligence from a variety of provers seamlessly and better manage costs. This will be an important part of that.
Aaron Levie
Stripe + OpenRouter coming together is huge. For AI to diffuse more broadly, developers and enterprises will want ways of being able to mix and match intelligence from a variety of provers seamlessly and better manage costs. This will be an important part of that.
Aaron Levie
As we’re seeing in case study after case study, it turns out that the amount of value that can be created between the AI model and the ultimate end-user workflow is far larger than many people assumed or realized. Model capability is obviously doing a lot of heavy lifting in agentic products, but there’s still a lot more work to diffuse AI into the enterprise. 1. Getting agents to work well (and alongside people) in mission critical workflows tends to need to be represented differently depending on the business process. Sometimes it’s a chat experience. Other times it’s a background agent running in a deterministic workflow. And dozens of other variants. This is a mix of needing a harness that’s tuned to specific domains of work, but also making it show up in the right product experience. 2. Different workflows connect into entirely different enterprise systems and need access to very different data. Working with that data -whether it’s life sciences, financial, legal, etc.- requires contextual approaches, understanding of the data, having the right user experience for data interaction, and more. 3. The need for domain-specific change management remains critical in most verticals. The way you talk and implement technology at a bank is very different from a law firm. Having the right talent with a singular mission ends up being extremely useful for something as complex as process reengineering. 4. The ability to work with a variety of models means you can tune the workflows to different cost and performance levels. And you can eventually post train models for specific tasks to tailor the outcomes and eke out gains that aren’t coming otherwise in frontier models. 5. Evals! AI is basically not useful if it can’t be evaluated. Domain-specific evals that let you dramatically improve the performance of your harness for specific workflows just has a crazy long tail given how many tasks there are in the economy. Nearly impossible for one system to be tuned for all of them. 6. Lots of verticals and domains require pricing models that reflect relevant abstractions on top of tokens alone. The ability to price in ways that work for your industry’s consumption model ends up mattering in a variety of spaces. This just touches on some of the things that go into the applied AI layer. But it all adds up to being a huge surface area for being able to sustainably innovate and differentiate.
Aaron Levie
When you hear that data is the new oil, this is ultimately what that looks like. AI has such a thirst for data that we’re entering an era where it’s valuable almost in any form. In a world of AI, information actually belongs as an asset on the balance sheet. This data sale is just one example of the implications of what that will look like. Ultimately, the way that companies manage and mine their organization’s intelligence will be one of the deciding factors for competitiveness and value creation in the future.
Aaron Levie
AI spend is nowhere close to hitting any walls. Obviously this data is weighted toward engineering-centric companies, with the top 1% spending $7,500/mo and top 10% spending $660/mo per employee on AI. But the general trend of continued exponential growth is occurring throughout all firm types. And what the top 10% is doing today may easily be what the top 50% is doing in 3 years from now, at least in terms of token volume if not in dollars. This is why there’s still so much opportunity right now in the diffusion of AI in the enterprise. Especially as token costs come down, we will throw larger portions of work at agents. They’ll be scanning all our code for security issues, testing all of our software, writing coding for much larger projects, processing nearly all data, and much more. So this trend has no end in sight.
Aaron Levie
Pretty remarkable what’s happening with open weights AI right now. We’re seeing models achieve SOTA results on specific tasks, and getting close to frontier on some areas of coding and other domains. The more that open weights is able to maintain only a marginal gap from the…
中文: 目前开放式重量人工智能正在发生的事情相当惊人。我们看到模型在特定任务上实现SOTA结果,并在编程和其他领域接近前沿。 开量权重越能保持与......的微弱差距
Aaron Levie
The fact that open weights models are being discussed credibly at this level of capability should be a huge update for many. The implications of open models getting to frontier performance ensures that you can always have sovereign AI, have the ability to post train for your…
中文: 开放权重模型在这种能力水平上正在被可信地讨论,这对许多人来说应该是一次巨大的更新。 开放模型达到前沿性能的影响,可确保您始终拥有自主的人工智能,能够为您的...发布列车
Aaron Levie
This whole Fable export control situation is actually net positive to regulation discourse. It’s an early peek into what AI regulation would end up looking like at scale when enacted at the model layer instead of the specific application of the AI. The government would have sole…
中文: 整个“寓言”出口管制形势实际上对监管讨论持正面态度。这是对人工智能监管在模型层实施时最终会大规模实施的情况的早期观察,而不是人工智能的具体应用。 政府将拥有唯一的......
Aaron Levie
Token costs are becoming one of the hottest topics for any enterprise I talk with right now. It’s very bullish for AI in general because it means these systems are being used at a scale that wasn’t contemplated before. It also gives way to another form of differentiation that…
中文: 代币成本正成为我目前所谈论的任何企业最热门的话题之一。这总体上对人工智能非常看好,因为这意味着这些系统正以前所未有的规模被使用。 它还让位于另一种形式的分化......
Aaron Levie
Again, maybe counterintuitive, but in the majority of conversations I have with CIOs, CTOs, and CEOs in large enterprises, they are either growing due to AI (in new job functions like FDEs, engineering, etc.) or at a minimum reinvesting efficiency savings back into the business…
中文: 同样,可能具有反直觉性,但在与大型企业的首席信息官、首席技术官和首席执行官们进行的大多数对话中,它们要么因人工智能(如FDE、工程等新工作职能)而增长,要么至少将效率节约重新投入到企业中......
Aaron Levie
Another week on the road meeting with a couple dozen IT and AI leaders from large enterprises across banking, media, retail, healthcare, consulting, tech, and sports, to discuss agents in the enterprise. Some quick takeaways: * Clear that we’re moving from chat era of AI to…
中文: 又一周与来自银行、媒体、零售、医疗、咨询、科技和体育等大型企业的数十名IT和人工智能行业负责人举行客场会议,讨论企业中的代理商。 一些快速的外卖: * 明确我们正从人工智能的聊天时代走向......
Aaron Levie
Box just launched its plugin within Codex, which means you can take any content within Box and automate workflows around it using the power of a coding agent. Here's a quick example of processing earnings call documents to extract structured data at scale, which you could then… https://twitter.com/levie/status/2037383187245273477/video/1
中文: Box 刚刚在 Codex 中推出了其插件,这意味着你可以将任何内容都采用 Box,并利用编码代理的功能实现其工作流程的自动化。 以下是处理收益电话文件以大规模提取结构化数据的快速示例,您可以将其处理到此过程......
media 0 共 2 项
🎬
视频