Asuka Zheng🦭
中文: 没有上限,我们收到了ROPE。
Asuka Zheng🦭
no it's yukio mishima.
中文: 不,这是yukio mishima。
Asuka Zheng🦭
can’t believe it’s October already. 2025 is ending soon.
中文: 简直不敢相信已经是十月了。 2025即将结束。
Asuka Zheng🦭
中文: 🦭 参加了一场海滩派对
Asuka Zheng🦭
💀🦭zelda and overwatch changed my dark life. https://twitter.com/VoidAsuka/status/2104109832227955036/photo/1
Asuka Zheng🦭
He jurgen on their huber til we schmid.
中文: 他淺在他们的丈夫身上,直到我们小朂密。
Asuka Zheng🦭
i said i'd study Miles and so here is my first little pr, using H100s on @sfcompute's givemeanode, now Autoresearch. thank you my friend @evanjconrad <3🦭 Autoresearch is the easiest GPU platform i've used! i used to write and maintain a lot of scripts to use rented compute efficiently. with Autoresearch, my coding agent could create nodes, run the Miles container, inspect failures, retrieve results, and shut down the H100 node when I no longer needed it - all in the same workflow as editing code. It saved me a lot of manual node setup, file management (and money!) the Diffusers version Miles pins could override the deterministic setting and wrapped FA3 in a custom op without a registered backward. my patch calls FA3's public autograd interface and preserves the setting through dispatch. on H100s, the attention tests matched numerical references and produced bitwise-identical outputs and gradients across 20 fixed-input runs per configuration, including 2- and 4-GPU Ulysses. i also added a tiny Wan FSDP2 training test comparing SP=1 and SP=2. GPU charges for this validation came to about $8, including setup and retries(very efficient!). draft PR: https://github.com/radixark/miles_diffusion/pull/2
中文: 我说我学了迈尔斯,所以这是我的第一个小PR,使用H100s来使用@sfcompute的赠与度,现在为Autoresearch。谢谢我的朋友@evanjconrad Autoresearch 是我用过的最简单的 GPU 平台!我过去常常编写和维护大量脚本,以便高效使用租用的计算。通过Autoresearch,我的编码代理可以创建节点、运行里程容器、检查故障、取回结果,并在不再需要时关闭H100节点——这些节点的工作流程与编辑代码完全相同。它为我节省了大量手动节点设置、文件管理(以及资金!) Diffusers 版本的 Miles 引脚可以覆盖确定性设置,并将 FA3 包裹在自定义操作中,无需向后注册。我的补丁将 FA3 的公有自动渐回界面称为,并通过发送来保留设置。 在H100上,注意力测试与数值参考值匹配,并在20个固定输入运行(包括2和4-GPU Ulyses)上生成了相同的输出和梯度。i还添加了一个微小的Wan FSDP2训练测试,比较了SP=1和SP=2。 此验证的GPU费用约为8美元,包括安装和重置(非常高效!)。 PR草案:
Asuka Zheng🦭
awesome work and awesome tech report! ever since minimax h3, there’s been so much acceleration work around omni models lately - everyone wants to make them real-time. it’s really interesting to watch these approaches converge. sparse and low-bit attention feel especially complementary: sparse attention reduces how many interactions you compute, while low-bit attention reduces the cost of each interaction. real-time omni models may end up being one of the strongest forcing functions for attention and inference efficiency.
中文: 出色的工作和出色的技术报告! 从 minimax h3 起,最近就已对全模型进行了如此多的加速处理——每个人都希望让它们成为实时的。 观察这些方法的融合确实很有趣。稀疏且低比特的注意力尤其令人感到互补:稀疏的注意力会减少你计算的交互量,而低位注意力则能降低每次交互的成本。 实时全模式可能最终成为注意力和推理效率最强的强制函数之一。
Asuka Zheng🦭
RT @paradigmainc: Introducing Limite 1B - Violetto. A model for high-frequency mathematical intelligence. https://twitter.com/paradigmainc/status/2102101878305698058/photo/1
中文: RT @paradigmainc:介绍 Limite 1B - Violetto。 高频数学智能模型。
Asuka Zheng🦭
2030: Your AI lawyer thinks you’re guilty, so the prosecutor subpoenas its activations.
中文: 2030年:你的人工智能律师认为你有罪,因此检察官会发出传票,要求激活。
Asuka Zheng🦭
&gt;Their entire hacking cost less than $3,000 in tokens. &gt;OpenAI awarded them $6,500.
Asuka Zheng🦭
Following MiMo-V2.6’s live RL run. Useful to see learning curves alongside sampling stats, queues, costs, and restarts! The2B tokens per step is a substantial workload. What I’d like to understand is the compute allocation: at a fixed total budget, what produces the most improvement on unseen tasks? A few questions this run brings up: - More prompts, more rollouts, or better feedback? The dashboard shows ~1,568 prompts × 16 rollouts per batch. Covering more tasks and exploring more trajectories per task serve different purposes. Grader compute adds another choice: when is it worth generating fewer trajectories and spending more on evaluating them? The best allocation may shift as the policy improves. - What does agentic in-group credit assignment actually do? Test cases and rubrics provide different kinds of feedback. I’m curious how the grader compares attempts within a group, and whether it assigns credit to whole trajectories, stages, or individual actions. The ablation I’d like to see is whether additional grading cost pays for itself through more efficient learning. - Critic/RLOO choices also affect scheduling. RLOO uses sibling rollout returns as a baseline, a learned value function predicts return from the current state. If advantage computation requires the full group, slow trajectories can delay it even in an asynchronous pipeline. A value-based estimator that permits independent trajectory processing could reduce that dependency, but adds critic inference, training, and estimation costs. Retaining an RLOO component may retain the group dependency too. This is a design question, not a claim about MiMo’s implementation. The dashboard’s critic/* metric names alone don’t establish that it uses a learned critic. - Throughput needs to be read alongside policy lag. More trajectories per hour don’t necessarily mean faster learning. By the time a long-running task finishes, its behavior policy may be several versions behind. How the system handles that gap - and avoids systematically favoring short tasks - matters alongside raw throughput. The comparison I’d most like to see: time and total cost to reach the same held-out performance, accounting for rollouts, grading, environments, and any critic. Appreciate the MiMo team making the run visible. There’s a lot to learn from the intermediate behavior of a training system. https://mimo.xiaomi.com/rl/ Le Critique: https://arxiv.org/abs/2608.16739
Asuka Zheng🦭
my bet: future agent harnesses will be orchestrations of models, with trained smaller models handling routing, context isolation and memory updates. models deciding what other models get to see, remember and spend compute on.
中文: 我的赌注: 未来的代理工具将采用模型编排,配备经过训练的小型模型,处理路由、上下文隔离和内存更新。 模型决定其他模型能够查看、记住并花费计算。
Asuka Zheng🦭
RT @maarcoofdezz: la vida antes de este tweet: https://twitter.com/maarcoofdezz/status/2097332381912879337/video/1
🎬
视频
Asuka Zheng🦭
I heard it's the best post-training framework right now, with really good engineering taste. Gotta study and try it out carefully this week! 👀
中文: 我听说这是目前最好的训练后框架,具有非常优秀的工程品味。 本周一定要学习并仔细尝试!👀
Asuka Zheng🦭
RT @maxxrubin_: So Astra is able to identify sounds from mel spectrograms zero-shot. I don't think we've scratched the surface of what this model can do (and this is light reasoning btw) https://twitter.com/maxxrubin_/status/2096892510241268094/video/1
media 0 共 4 项 media 1
🎬
视频
🎬
视频
Asuka Zheng🦭
I think interpretability might not be a problem to be solved, but rather an engineering tool we can use. It might not be scalable, and we may need to patch things a lot with it, but it will still be useful.
中文: 我认为可解释性可能不是需要解决的问题,而是我们可以使用的工程工具。它可能无法扩展,我们可能需要用它大量修补,但它仍然会很有用。
Asuka Zheng🦭
so again, most tasks can be considered coding tasks, and the results are natively programmable.
中文: 因此,大多数任务都可以被视为编码任务,结果是原生可编程的。
Asuka Zheng🦭
rsi 😂
Asuka Zheng🦭
RT @yuntiandeng: Wow, this brought back a project Woojeong Kim, @srush_nlp, and I worked on in 2024. A strong model read a task description, turned it into "ability vectors", and passed them to a weak model. Sasha even came up with a great name for it: Strong-to-Weak Ability Transfer (SWAT) 😆 We never published it, but the idea eventually developed into ProgramAsWeights, which compiles task descriptions into small reusable neural programs that run locally. Very cool to see @aimalysheva and the Mostik team make a related idea work so well at this scale. https://programasweights.com/
Asuka Zheng🦭
this is actually huge. sota models could directly provide hidden states to smaller models, allowing a 4B parameter model capable of running on edge devices to perform decoding. it's a new way to do fast and cheap inference. shut out to people who first made it. will agent talk in latent space in the future? probably.
中文: 这实际上非常大。Sota 模型可以直接为较小的模型提供隐藏状态,从而实现能够在边缘设备上运行的 4B 参数模型进行解码。 这是一种快速且廉价推断的新方法。 对最初制作它的人来说很封闭。 代理人将来会在潜在空间里说话吗?可能。
Asuka Zheng🦭
I'm really curious about how Jürgen sees the world. The guy is standing so high up and looking so far ahead. The more you read his work, the more you realize it. A kid, an artist, and a master.
Asuka Zheng🦭
We definitely live in a simulation. It also reminds me of @MingchenZhuge's work on neural computers: the whole interactive computer terminal as a world model.
中文: 我们绝对生活在一个模拟中。 这也让我想起了@MingchenZhuge在神经计算机方面的工作:整个交互式计算机终端作为世界模型。
Asuka Zheng🦭
RT @yuzu_4ever: @UpdateLiveware i really liked summer 2024 but the last two have been a bit rough
中文: RT @yuzu_4ever:@UpdateLiveware 我很喜欢2024年夏季,但最近两款有点粗糙
Asuka Zheng🦭
i guess we will never see those use cases happening everywhere if sota video models remain entirely closed-source. 🥺(seedance 2.0 was released in feb, h3 was released in aug.) with the release of fable 5.1 today, i feel an urgent push to use as many tokens as possible to learn and create - because i'm low-key scared that one day we won't be able to access sota llm models so easily. (also, distillation is getting harder) this is also why open source has to win.
中文: 如果 sota 视频模型完全封闭,我们绝不会看到这些用例随处可见。🥺(《种子2.0》于fef发布,h3于8月发布。) 随着今天Fabre 5.1的发布,我迫切需要使用尽可能多的代币来学习和创建——因为我低调地担心,总有一天我们无法如此轻松地访问sota llm模型。(此外,蒸馏越来越难)这也是开源必须取胜的原因。
Asuka Zheng🦭
more and more to come!🤗
中文: 来得越来越多!EE0
Asuka Zheng🦭
RT @wlsaidhi: My guy, are you seriously as an 8 billion “commercial entity” telling an open source project that distillation is not a game of compute and data??? Meanwhile you guys take an open source model, post train it on user data, and keep it closed source behind API while you complain about our preview open weight release using disingenuous graphs. Have some self awareness man
中文: RT @wlsaidhi:我这家伙,你作为一个80亿“商业实体”,真的在向开源项目发帖,说蒸馏不是计算和数据游戏吗? 与此同时,你们采用开源模型,在用户数据上发布训练,并将其保持在API的闭源位置,同时通过不诚实的图表投诉我们的预览开放权重发布。 有自我觉知的人
Asuka Zheng🦭
try post train H3 with sglang miles!
Asuka Zheng🦭
RT @wlsaidhi: My guy, are you seriously as an 8 billion “commercial entity” telling an open source project that distillation is not a game of compute and data??? Meanwhile you guys take an open source model, post train it on user data, and keep it closed source behind API while you complain about our preview open weight release using disingenuous graphs. Have some self awareness man
中文: RT @wlsaidhi:我这家伙,你作为一个80亿“商业实体”,真的在向开源项目发帖,说蒸馏不是计算和数据游戏吗? 与此同时,你们采用开源模型,在用户数据上发布训练,并将其保持在API的闭源位置,同时通过不诚实的图表投诉我们的预览开放权重发布。 有自我觉知的人
Asuka Zheng🦭
open-source state-of-the-art models unlock the potential for everyone and every team to create something historical and push the frontiers of model capability. so happy to see what MiniMax-H3 brings to the world!🥰
Asuka Zheng🦭
open weight and open recipe on MiniMax FastH3 from the OG of video acceleration lab.👩‍🍳
中文: 视频加速实验室OG的MiniMax FastH3上的开放式重量和开放式食谱。EE0 🍳
Asuka Zheng🦭
i'll be sharing how to fine-tune and host H3 tonight in SF! i'll also walk through inference speed benchmarks across different hardware setups and recipes.
中文: 今晚在旧金山,我将分享如何进行微调和主持H3!我还将在不同硬件设置和配方中浏览推理速度基准。
Asuka Zheng🦭
RT @antirez: Btw the Redis story repeats itself: I'm working at DwarfStar for free for the community and because I enjoy it. But I'm receiving criticisms, since people are worried that this will break their AI-richness plans. I want to say to everybody thinking that I should stop that each time you tell me this, I'll double down my efforts towards a completely no profit engine for local inference. Better to shut up basically.
中文: RT @antirez:不过,Redis 的故事会重演:我为社区免费在 DwarfStar 工作,因为我乐在于此。但我正受到批评,因为人们担心这会破坏他们丰富的人工智能计划。我想对所有想让我停下来的人说,每次你告诉我这件事,我都会加倍努力,为本地推断完全不盈利。最好基本上闭嘴。
Asuka Zheng🦭
complier for inference and programmable llm
中文: 推理更为 可编程的 lm
Asuka Zheng🦭
We'll soon share how we post-train MiniMax H3 using the SGLang framework!
Asuka Zheng🦭
Most tasks can be framed as coding tasks, since coding is just a layer of abstraction. Once any process has a standard SOP, it becomes programmable.
Asuka Zheng🦭
RT @cwolferesearch: Getting ready to publish my complete guide to RL for LLMs tomorrow morning. Although the post contains many of my own thoughts / learnings, it is also a synthesis of so many great resources that have been published over the years: - The RLHF Book (https://t.co/n3aUxqWtgS) by @natolambert - Reinforcement Learning by Richard S. Sutton and Andrew G. Barto - Spinning Up in Deep RL (https://t.co/EYcOOllvIy) from OpenAI - Build an LLM (https://t.co/6HfXuofVwq) and Reasoning Model (https://t.co/l5uS247A04) from Scratch by @rasbt - Various notes (https://t.co/hpsyAKOnRv) and papers (TRPO, PPO, etc.) from John Schulman - Policy Gradient Algorithms (https://t.co/caNLakSjo5) by Lilian Weng - A Vision Researcher’s Guide to RL (https://t.co/25awrjj5Bd) by @YugeTen - From REINFORCE to Dr. GRPO (https://t.co/M23e3Fl3VZ) by @qingfeng_lan - Async GRPO in the Wild (https://t.co/A5qeYMwcy4) by @yumo_xu - Open RL infrastructure like TRL (https://t.co/TGrrnJ574t) and OpenInstruct (https://t.co/L9pObiUb1c) I highly recommend reading all of them. They’ve truly helped me to learn so much.
Asuka Zheng🦭
huge. very lucky to see the live demo - genuinely shocked. the speed changes the unit economics, and both unlock so many possibilities for video generation. this kind of work will also benefit world models and robotics models down the line: once model capability is good enough to use, inference speed and cost start to matter a lot. feels like a good direction for early bets: accelerating video + multimodal inference, possibly even to near real-time.
Asuka Zheng🦭
white-pilled
Asuka Zheng🦭
thanks to @NicoleSHsing @JentseHuang for making it happen🫶
Asuka Zheng🦭
Asuka Zheng🦭
need it
Asuka Zheng🦭
Lowkey, I think we need a harness benchmark to evaluate the performance of different harness frameworks - using the same model or the same set of heterogeneous models - on performance, cost, and latency. (Yes, I think we’ve reached an era where latency has become a major issue now.) There are still so many things we can do, and so many variables to consider.
Asuka Zheng🦭
my afternoon tea activity is reading the best papers from last week in my garden. &lt;3 Sparse Attention from @MiniMax_AI and From AGI to ASI from @GoogleDeepMind get some sun, folks - the tax here is so high. https://twitter.com/VoidAsuka/status/2066674079848157200/photo/1
中文: 我的下午茶活动是阅读上周在花园里最好的论文。 &lt;3 @MiniMax_AI 和 @GoogleDeepMind 的 AGI 到 ASI 的 Sparse 注意力 要晒太阳,各位——这里的税分高。