RT @hrkrshnn: We just released apex-flash-1, an open-weights model we post-trained for cybersecurity. Fire up your GPUs and run it!
We've post-trained models from 27B parameters all the way to 1T+ parameters. The unique nature of cybersecurity is that if you can build a system that finds zero-days at scale, you've built a machine that can print money.
You can test it on public bug bounties and profit from doing that. We've earned a million dollars in bounties across various programs, and we're currently #1 on the HackerOne US business leaderboard for 2026! Companies like Apple, Anthropic, Datadog, Ripple, Coinbase have paid us well for our disclosures.
Last year was all about harness engineering and riding the wave of models getting more intelligent over time. We built harnesses for offensive and defensive security work that process trillions of tokens a month. Yes, trillions!
With all that work, it's clear to us what the future of security looks like.
Hacking was once a bespoke skill, similar to hardcore software engineering. Software engineering is currently moving from offices to factories. Security, too, will go through the same transition.
Last month, Anthropic released a report on how they caught different threat actors abusing Claude. One of them stood out to me. A team in China built an automated exploit factory. A factory that autonomously grinds through vulnerabilities in targets and exploits them. Their targets included security products, network appliances, and government organizations.
We've spent enough time building systems adjacent to that to know what it takes, and the shocking part is the economics. How little it costs to breach companies that you and I rely on.
All this to say, security is largely becoming an economics problem. Reasonably capable models and harnesses, with enough compute, can find ways to steal your data, damage your reputation, and sometimes even take your money. And the cost to achieve that is dropping every day.
When that happens, the right question to ask is: what's the cost, and what's the lowest cost to achieve an outcome reliably? You want to find a stack with the best combination of capability and cost. This is what people call 'Pareto-optimal'.
Once you know that, it's all about scaling. Scaling to trillions of tokens a month, then a week, then a day. Eventually, trillions of tokens a second.
To do this, you need to own and control the entire stack: the models, the harnesses, the context, and even the flow of tokens. If you do the math here, the economics start looking insane.
What does that look like?
1. Building evaluations on cyber tasks that you and your customers care about.
2. Building a data pipeline of unique real-world data with signal attached to it.
3. Building a harness that can self-improve for each specific organization or customer.
4. Building a post-training loop that can take those learnings and improve the models. Models that are better, faster, and cheaper.
5. Owning your inference pipelines and controlling the flow of tokens so you're maximizing the value produced in every GPU cycle.
On evals: if you're a company building AI products for customers, you have to build your own. There's tremendous alpha in having internal evals corresponding to real work. In our case, these are evals for offensive and defensive security work.
A lot of public evals are bad for measuring the work you actually care about. Public evals are typically from academics or data companies.
Academics have limited budgets and limited access to proprietary data that measures real economic work. Data companies build evals and also sell you corresponding data (RL gyms, traces, pre-training data, etc.) that helps your next model “juice up” its score.
This is the dirty secret, and also why models can do extremely well on public evals but extremely poorly on things you care about. This is also why there's tremendous alpha in private evals: you know which models are the best for your use cases.
In cyber, some of the popular public evals are ExploitGym and ExploitBench. We think these evals do not correspond to the real security work our customers need. A typical challenge in ExploitBench is to take Chrome's javascript engine V8 with a known patch and find a way to exploit the vulnerability.
That is far, far removed from the security issues in everyday applications that you and I use, and applications built by our customers.
We could've benchmaxxed on these evals, but we didn't. And you should be wary of people using these evals to advertise their models and harnesses. Measure things yourself and see how they apply to your use cases.
The future: Opus is a model that many people found to be a reliable workhorse. For me, Opus 4.1 was the first model that was functional and could get work done.
These days, the flash models from Qwen, DeepSeek, MiMo and GLM can be categorized as workhorses too. That, combined with owning the post-training loop means the economics start looking insane. You can get frontier performance in specific domains for a small fraction of the price of Opus.
With the right data pipelines and post-training loop, you now have a durable strategy to keep improving and stay state-of-the-art in your vertical.
That's the bet behind apex-flash-1, our post-train of glm-5.3-flash, and we're already preparing for longer training runs.