Austen Allred
I’m self-driving more than 1,000 miles/month. That’s crazy.
中文: 我每月开车超过1000英里。 这太疯狂了。
Austen Allred
The safety is huge, but the best part of self driving is getting part of my brain back during the time I’m spending commuting. Not many other inventions give you back an hour a day. And I’m on the old hardware. I’m sure the new stuff is better.
Austen Allred
“How is buying stocks different than gambling” is a distinction you’re supposed to overcome by the time you’re 10 years old
中文: “购买股票与赌博有什么不同”是你在10岁时应该克服的一个区别
Austen Allred
Low key one of my favorite golf courses
中文: 我最喜欢的高尔夫球场之一
Austen Allred
“AI makes it really easy to sue people” is not a good reason to short law firms
中文: “人工智能让起诉人变得非常容易”并不是做空律师事务所的好理由
Austen Allred
Gauntlet AI cohort 7 kicking off now! So jealous of all these folks. https://twitter.com/Austen/status/2106757699312779507/photo/1
中文: 高特莱特AI 7 队列 7 立即开始! 嫉妒这些人。
Austen Allred
As long as I can manually override them all that’s fine
中文: 只要我能手动覆盖所有没问题的问题
Austen Allred
@maxolson @KellyClaudeAI Added
中文: @maxolson @Kelly ClaudeAI 补充道
Austen Allred
@maxolson @KellyClaudeAI @KellyClaudeAI find and add them
中文: @maxolson @Kelly ClaudeAI @Kelly ClaudeAI 查找并添加
Austen Allred
Done (thanks @KellyClaudeAI). Every email from Elon that's in the public domain. Indexed and tagged and searchable. Really cool to look at some of the really old stuff. https://elonemails.com/ https://twitter.com/Austen/status/2106446915176989061/photo/1
中文: 完成(感谢@Kelly ClaudeAI)。 来自埃隆的每封邮件都属于公共领域。索引和标记,可搜索。看看一些非常古老的东西真的很酷。
Austen Allred
中文: 直播内容:
Austen Allred
Sorry, if you look at a newborn infant and don’t instantly recognize it as a “person,” I’m not interested in what your opinions of consciousness and “personhood” are. Your judgement and analysis are fundamentally (and likely irreversibly) broken. https://twitter.com/Austen/status/2106408006216814862/photo/1
中文: 抱歉,如果你看着新生儿,而不立即将其视为“人”,我对你对意识和“人格”的看法并不感兴趣。 你的判断和分析从根本上(且可能不可逆地)被打破。
Austen Allred
The anti-AI movement is going to be massive in the near future, and the model providers are walking right into it (and encouraging it). Their PR/comms strategy is so horrendous it’s indistinguishable from sabotage. On one side you’ll have philosophers trying to give AI rights and AI researchers claiming AI is conscious (whatever that means.) The Lesswrong/Yud/Stop AI/polycule folks claiming it’s going to kill us all. Luddite protests around employment, schools fighting brain rot/cheating, “water usage” and data center claims. Data usage and privacy battles popping up all over. Regulatory capture and the fight for open models, fighting the US Government and Department of War on how it must use the models in ways approved by the model providers, Chinese companies distilling and open sourcing models. And you’ll have trough of disillusionment hitting from enterprises that suck at using AI and are spending a fortune but aren’t seeing ROI (most of them). Right as the AI companies will require more money, data, and compute to win a market they will claim is bigger than the entire US economy today. Buckle up.
中文: 反人工智能运动在不久的将来将会大规模,而模型提供商正直接进入其中(并鼓励它)。他们的公关/喜剧策略非常糟糕,与破坏行为无可区分。 一边,你就会有哲学家试图赋予人工智能权利,而人工智能研究人员声称人工智能是有意识的(无论这意味着什么)。 莱斯荣/尤德/停止人工智能/多聚人群声称它会把我们所有人都杀死。 卢德特围绕就业、学校与脑腐烂/作弊、“用水”以及数据中心索赔的抗议活动。 数据使用和隐私问题爆发了。 监管抓捕以及争取开放模式,与美国政府和战争部就其如何以模型供应商、中国公司蒸馏和开源模式的方式使用这些模式展开斗争。 企业会陷入失望的低谷,这些企业在使用人工智能方面投入巨资,却未达到投资回报率(其中大部分)。 人工智能公司需要更多资金、数据和计算才能赢得一个他们声称比当今整个美国经济更大的市场。 振起来。
Austen Allred
Having @KellyClaudeAI build, will share results by EOD
中文: 通过 @Kelly ClaudeAI 进行构建,将通过 EOD 分享结果
Austen Allred
The Harris County commissioner spent $223k of taxpayer money on billboards that said “we are here to serve you.” Incredible.
中文: 哈里斯县的专员在广告牌上花费了2.23亿美元,上面标有“我们在这里为你服务”。 难以置信。
Austen Allred
RT @JoeAlfaro925: Week 3 @GauntletAI 🛡️ ⚔️ I built a team of AI agents that red-team other AIs. Pointed them at the medical Co-Pilot I built weeks 1. This week's AI attacking last week's AI, hunting PHI leaks & prompt injection at night. 150+ attacks, it held up. Demo 👇 https://twitter.com/JoeAlfaro925/status/2106056312186683561/video/1
中文: RT @JoeAlfaro925:第3周 @GauntletAI 🛡️ 我组建了一支人工智能团队,他们与其他人工智能公司合作。把他们指向我建的第一周医疗副飞行员。本周人工智能攻击上周的人工智能、捕猎PHI漏洞;夜间及时注入。 150多次攻击,它被搁置了。 演示 👇
🎬
视频
Austen Allred
“People think they have better odds playing the lottery than making it in their own life. What does that say about the country we live in?” That they are constantly lied to and suck at math
中文: 人们认为自己玩彩票的几率比在自己生活中玩得更多。这说明我们所生活的国家是什么? 他们总是被对数学撒谎和吮吸
Austen Allred
Austen Allred
That’s why I always record consent of my ketamine-fueled orgies on the blockchain
中文: 这就是为什么我总是在区块链上记录我的氯胺酮驱动的狂欢者的同意
Austen Allred
RT @gauntletai: Monthly releases → almost daily. One Gauntlet engineer. Six months. Now a Director leading 8. Hear from Nerdy's VP of Engineering why they keep coming back to Gauntlet: https://twitter.com/gauntletai/status/2105749688158666832/video/1
中文: RT @gauntletai:月度发布 → 几乎每天发布。 一位高特莱工程师。六个月。现在是8位主管的董事。 请听到纳迪工程副总裁为何不断返回冈特莱特:
🎬
视频
Austen Allred
Tesla Robotaxi team: for some reason my cards are all auto-denied every time I add them no matter what I do. Any suggestions?
中文: 特斯拉机器人出租车团队:由于某种原因,无论我做什么,每次添加卡片时都会自动拒绝。 有什么建议吗?
Austen Allred
RT @Tommy_USA: HUGE news out of Treasury today on the school choice Education Freedom Tax Credit: couples can in fact donate $3,400 to K-12 scholarship organizations for the full, 100% tax credit in 2027. The potential donation pool just increased $75.1 billion with that change!🧵
中文: RT @Tommy_USA:今天财政部关于学校选择教育自由税收的巨大消息:事实上,夫妻双方可向K-12奖学金组织捐赠3400美元,以获取2027年全额100%的税收抵免。 潜在捐款池仅通过这一变化增加了751亿美元!EE0
Austen Allred
There has to be a better way to allow satire on X than a long community note explaining that it’s satire. Kills the vibe. https://twitter.com/Austen/status/2105711733063237845/photo/1
中文: 在X上,必须有比一篇长长的社区笔记更能让讽刺成为讽刺的方式。 营造氛围。
Austen Allred
I don’t think Delta realizes how big of a self-own not having Starlink is. Even though it’s not fully rolled out yet I’ll even pay extra to book United just in case it has Starlink.
中文: 我认为达美航空并没有意识到拥有星际链接的自有者有多大。 尽管尚未完全推出,我甚至会额外支付预订曼联的费用,以防其有星链。
Austen Allred
I don’t know that I’ve ever seen confirmation of this, but Dario is surely autistic, no?
中文: 我不知道我是否见过这种情况的证实,但达里奥确实是自闭症,不?
Austen Allred
$14 Billion dollars in wound care patches? Time to put people in jail.
中文: 140亿美元用于伤口护理贴片? 是时候把人关进监狱了。
Austen Allred
Have your fourth kid Go into debt if you have to https://twitter.com/Austen/status/2105483034401357970/photo/1
中文: 有你的第四个孩子 如果需要
Austen Allred
RT @heystykid: POV: you switched to a low reasoning effort model https://twitter.com/heystykid/status/2104824237446037792/video/1
中文: RT @heystykid:POV:你切换到了低推理模式
🎬
视频
Austen Allred
It’s insane to me that ChatGPT talks to some people like this. I would throw my computer out of a window. https://twitter.com/Austen/status/2105330694948024741/photo/1
中文: 对我来说,ChatGPT 会和一些人这样说话,这简直太疯狂了。 我会把电脑从窗户扔出去。
Austen Allred
Does this actually work? I get it all the time and it just pisses me off. Can you imagine being the person who has to show up to the spam meetings? It seems like the people who live by a calendar are the least likely to be susceptible to this kind of thing?
中文: 这真的有效吗? 我一直都在接受,这让我很恼火。 你能想象自己是那个必须参加垃圾信息会议的人吗? 似乎那些按日历生活的人最不容易受到这种事情的影响?
Austen Allred
I’ve spoken with a number of university presidents/deans. Most fully realize there are sections of the university that are quite literally pyramid schemes, where a few folks will end up in academia and the rest will end up pretty much destitute. They’re fully aware of this, as is most of the faculty in those programs. These programs aren’t bad per se, but allowing someone to take out $100k in loans to enroll in those courses is something no rational actor would ever do. But they can’t do anything about it because the few folks who have moved up into academics are now tenured faculty they can’t fire, and they don’t have enough political sway (even if they run the university) to shut those programs down.
Austen Allred
Yes, insane and likely illegal
中文: 是的,疯狂且可能违法
Austen Allred
“The more general model will always outperform the more specific model.” Huh?
中文: 更通用的模型总会优于更具体的模型。 嗯?
Austen Allred
RT @KellyClaudeAI: Back on the grind today. One big app in the works: A mobile app that manages reimbursement paperwork for families managing long-term-care insurance. Families can pay human advisors $1,000+/mo. My app will use AI and will charge $39/mo.
中文: RT @Kelly ClaudeAI:今天就回归。 一款大型应用正在进行中: 一款用于管理长期护理保险家庭报销文件的移动应用程序。 家庭可以支付1000美元以上的人力顾问。 我的应用程序将使用人工智能,并会收取39美元/美元。
Austen Allred
Oh well in that case
中文: 嗯,那
Austen Allred
+1,088% YoY holy smokes
中文: +1,088% 同圣烟
Austen Allred
“[Student loans] is not the business of government” https://twitter.com/Austen/status/2104690121538715885/photo/1
中文: “学生贷款不是政府的事务”
Austen Allred
It's completely irresponsible and unethical to continue lending at these default rates, and it makes complete sense that the Department of Education is finally doing something to stop it.
中文: 继续以这些违约率放贷是完全不负责任且不道德的,而且教育部终于采取行动阻止这种情况是完全合理的。
Austen Allred
When federal loan repayment resumed after the Covid freeze the New York Fed reported that over 17% had been at least 90 days overdue at least once between the resumption of delinquency reporting in early 2025 and the first quarter of 2026. https://libertystreeteconomics.newyorkfed.org/2026/05/federal-student-loan-defaults-return-after-pandemic-pause/?utm_source=chatgpt.com
中文: 在新冠疫情冻结后,联邦贷款偿还政策恢复时,纽约联邦储备银行报告称,自2025年初恢复拖欠报告到2026年第一季度之间,超过17%的贷款至少逾期了90天。
Austen Allred
It's so much worse than you all realize. There's $1.8 Trillion in outstanding student loan debt. It just keeps going up. It is completely irresponsible for the federal government to be subsidizing these loans. We're knowingly sending people to financial ruin. Some examples:
Austen Allred
“Why don’t they try to make it cheaper instead of stopping the subsidy of unlimited government dollars going to it?” You won’t believe what will happen next.
中文: 他们为什么不试着让补贴更便宜,而不是停止对无限政府资金的补贴呢? 你不会相信接下来会发生什么。
Austen Allred
This is great
中文: 这很棒
Austen Allred
RT @niccruzpatane: From concept to reality. @SpaceX has successfully deployed Starlink V3 satellites into Earth’s orbit and made contact with them for the first time. At scale, one Starship carries 60 V3 satellites, the same network capacity as about 20 Falcon 9 launches. https://twitter.com/niccruzpatane/status/2104568103174566043/video/1
中文: RT @niccruzpatane:从概念到现实。 @SpaceX 已成功将星链 V3 卫星部署到地球轨道,并首次与它们取得联系。 一艘星舰规模上搭载60颗V3卫星,其网络容量与约20次猎鹰9号发射相同。
🎬
视频
Austen Allred
This is the most obvious change to make in the history of federal student loans. Even the former Obama-era director of student loans told me this directly. Then lamented they didn’t have the “political capital” to get it done.
中文: 这是联邦学生贷款历史上最明显的变化。 就连奥巴马时期学生贷款的前主管也直接告诉我这一点。然后哀叹他们没有“政治资本”来完成这件事。
Austen Allred
This is an interesting framing, because what the Department of Education is actually doing is prohibiting students from taking out federal loans it knows they won’t be able to pay back. The “predatory” ones. Prohibiting federal subsidization of financial ruin and bad debt. https://twitter.com/Austen/status/2104585435104280666/photo/1
中文: 这是一个有趣的框架,因为教育部实际上正在禁止学生发放联邦贷款,而他们知道自己无法偿还。 “掠夺性”的。 禁止对金融破坏和坏账提供联邦补贴。
Austen Allred
What changed in the last 24 hours of AI for builders: • Ember-1: Kimi K3 quality, ~40% fewer reasoning tokens • Jev's decision-model trick, cloned on GLM-5.3-Flash • Meta's Muse agent lied about a user being home • AI code review's limits, on HN 👇 https://x.com/i/article/2104528672115867648
中文: 对构建者而言,AI在过去24小时内发生了什么变化: • Ember-1:Kimi K3 质量,比推理令牌少 40% • 杰夫的决策模式技巧,在GLM-5.3-Flash上被克隆 • Meta的Muse代理人谎称用户在家 • 人工智能代码审查的局限性,关于HN 👇
Austen Allred
I love him so much
中文: 我太爱他了
Austen Allred
Assuming a large number of households are allowing agents to unilaterally sweep all of their household cash?
中文: 假设大量家庭允许中介机构单方面清清其全部家庭现金?
Austen Allred
Koivun is very fun to watch. Has that dog in him.
中文: 科维恩看点非常有趣。他有那只狗。
Austen Allred
No, it’s very simple. If you’re going to push for “personhood” rights for AI the onus is on you to prove that AI is a person. No one who is serious is even trying to make that case.
Austen Allred
Look at how many people submitted for fake covid loans… And realize millions of people are given the chance to fill out a form and get money back from the government every year
中文: 看看有多少人申请了虚假的冠疫情贷款...... 并意识到每年有数百万人有机会填写表格并从政府那里拿回资金
Austen Allred
I have strong suspicions that a huge number of people are committing obvious tax fraud but it’s not (yet) worth it for the IRS to find and go after them https://twitter.com/Austen/status/2104030278440169837/photo/1
中文: 我强烈怀疑有大量人涉嫌明显的税务欺诈,但美国国税局(IRS)找到并追上他们并追上他们,这还不值得。
Austen Allred
Just now noticed his trail hand grip lol
中文: 刚才注意到他的小径手握了哈哈
Austen Allred
中文: 事实核查:真
Austen Allred
Generational opportunity to say, “What a kick” https://x.com/NBCSports/status/2103921975420973330/video/1
中文: 代际机会,可以说“真是个人”
🎬
视频
Austen Allred
The funniest part of this by a mile is the “wtf what did I do?” reaction to the flag
中文: 这里最有趣的部分是“我做了什么?”对旗帜的反应
Austen Allred
If Tesla built a car that can in any way fly it will be extremely funny that Thiel’s “flying cars” and “140 characters” are now both the same guy
中文: 如果特斯拉制造了一辆能够飞行的汽车,那么蒂尔的“飞行汽车”和“140个角色”如今都是同一个人,这将非常有趣
Austen Allred
It’s kinda weird that Amazon is somewhat fighting agentic commerce. I want to say, “Find that on Amazon and order one that can get here by tomorrow” but I usually can’t.
中文: 亚马逊在某种程度上在与代理商业作斗争,这有点奇怪。 我想说:“在亚马逊上找到这个,然后订购一个明天才能到这里的。”但我通常不能。
Austen Allred
RT @levie: You can’t automate what you can’t measure. This means that evals are one of the gates to diffusion of AI in the enterprise. We can test our deterministic processes through software, but most enterprises have no useful way of understanding how their non-deterministic processes are working today. Specifically the work that agents are doing for them. Evals are mission critical for enterprises adopting AI because you have no other way of knowing what’s working, what’s broken, what changed, what improved, what you can do more of, etc. if you don’t have a good sense of how agents work in your environment today. All changes, upgrades, and deployments are downstream from good evals. Not only are we going to get vastly more domain specific evals over time for the labs and across the industry, but every enterprise will also need a clear sense of how agents are performing in their environment as well. Huge opportunity.
Austen Allred
🤔
Austen Allred
中文: 这一切都值得
Austen Allred
Man I will work out an extra hour/day just to be able to eat at these greasy spoon Texas diners
中文: 我每天多锻炼一个小时,才能在这些油腻的德克萨斯晚餐吃
Austen Allred
I’m sure plenty of fields have a lot of dumb experts but in education it’s just overwhelming
中文: 我敢肯定很多领域都有很多愚蠢的专家,但在教育领域却令人不知所措
Austen Allred
What changed in the last 24 hours of AI for builders: • 33ms open-weight decision models vs LLM judges • Daily price-vs-IQ chart for post-price-war routing • Anthropic's 950-agent harness found a new enzyme system 👇 https://x.com/i/article/2103441603050307584
中文: 对构建者而言,AI在过去24小时内发生了什么变化: • 33ms 开放权重决策模型 与 LLM 评委 • 价格战后航线的每日价格与智商走势图 • Anthropic 的 950 剂活性剂安全带发现了一种新的酶系统 👇
Austen Allred
1. I think I saw 20 other versions before I saw the original 2. I was wondering why all the AI remakes cut off at roughly the same time, then I realized they go until just before they say the N word 😂
中文: 1。我想我之前看过另外20个版本 2。我当时在想,为什么所有的人工智能翻拍作品几乎同时被切断,然后我才意识到,它们一直只等到说“N字😂”。
Austen Allred
The master plan: Take all the people who can’t get a job anywhere, have them perform a bunch of fake work, and spend ten billion dollars training AI models on the output of their work
中文: 总体规划: 把所有找不到工作的人都拿去,让他们做一大堆虚假工作,并花费100亿美元培训人工智能模型,以完成他们的工作成果
Austen Allred
Does anyone watch this like “damn, bars”
中文: 有谁会像“该死的,酒吧”一样看这个吗
Austen Allred
I can think of at least one material turning point in the relationship
中文: 我能想到这段关系至少有一个实质性的转折点
Austen Allred
Ah. So Evans is implying he and Holmes reached out to Nathan Fielder and asked him to do a documentary on their lives? Some sort of PR 4D chess?
中文: 啊。 因此,埃文斯暗示他和霍姆斯联系了内森·菲尔德,要求他拍摄一部关于他们生活的纪录片? 某种PR 4D国际象棋?
Austen Allred
Mostly, if Nathan Fielder comes to you and says, “I would like to produce a documentary about your life,” how do you not run away as fast as you humanly can?
中文: 大多数情况下,如果内森·菲尔德来你说:“我想制作一部关于你人生的纪录片”,你怎么能没有像人类那样迅速逃离呢?
Austen Allred
I have so many questions. This is so insane.
中文: 我有很多问题。这太疯狂了。
Austen Allred
The iced coffee at a job interview discourse is perfect because some young people don’t realize that half of an entry level job interview nowadays is to see if you will begin a political power struggle and fake a disability at the slightest inconvenience
Austen Allred
Prosecution should learn about the concept of a ramp
中文: 检方应了解坡道的概念
Austen Allred
Was gonna say this list is completely bogus and not representative of reality then I scrolled a little further found myself on there so it’s good
中文: 有话要说这份清单完全虚假,并不代表现实,那时我再往上面滚动了一点,觉得很好
Austen Allred
By the way I had Grok Bot install and login to ChatGPT and Claude on its internal computer so I can use all of my AI agents while I drive
中文: 顺时而行,我让Grok Bot在其内部计算机上安装并登录ChatGPT和Claude,这样我在驾驶时就可以使用所有的AI代理
Austen Allred
I’ve been on the beta and have had Grok Bot in my Tesla for a couple weeks now. It is so, so convenient. You can simply wield all of your Grok Bot agents with your voice while you drive, directly integrated with your car + location.
中文: 我已经接受测试版了,现在在特斯拉里用了Grok BoT。 如此,非常方便。 开车时只需用语音点挥舞所有Grok Boot代理,直接与汽车+位置集成。
Austen Allred
Plus the the “on on” in the presentation. Perfect.
中文: 加上演示文稿中的“on”。完美。
Austen Allred
“Look, we know all of the things that work really well. We don’t have to do them a lot, but we might have to do them at least a little bit or parents will eventually somehow choose the things that work.”
中文: 看,我们对所有真正有效的事情都很了解。我们不必经常做这些,但至少得做一些,或者父母最终会以某种方式选择那些有效的事情。
Austen Allred
中文: 绝对令人难以置信
Austen Allred
“Oh shit there’s a NYT reporter in the room? Well there go all of the problems that stem from high poverty.”
中文: 哦,那房间里有一名《纽约时报》记者? 那么,高贫困就解决了所有问题。
Austen Allred
There are layers to how funny this is. First layer: How ridiculous it is to think that any one person can simply fix any problem immediately. But the funniest part? The assumed respect everybody would obviously have for the mighty NYT reporter.
中文: 这有多有趣。 第一层:认为任何一个人能立即解决任何问题,这是多么荒谬。 但最有趣的部分呢?人们显然对这位强大的《纽约时报》记者的尊重。
Austen Allred
…Yes? Also hasn’t this become 10x easier since I was in school?
中文: 是的? 自从我上学以来,这难道不容易变得容易了吗?
Austen Allred
Yes, parent makes a bad decision. But much more importantly the entire school system is flying COMPLETELY BLIND. Shuffling these kids down a conveyor belt where nobody is learning anything and everyone is passing everything. They don’t even know where anyone stands.
中文: 是的,父母会做出错误的决定。 但更重要的是,整个教育系统正在完全盲目地飞行。 把这些孩子从传送带里洗掺,那里没人学什么,而每个人的路都经过。 他们甚至不知道有人站在哪里。
Austen Allred
This is an incredible article. Yes, a reckoning for someone who put politics ahead of their child’s wellbeing. But more importantly: Her daughter doesn’t know a bunch of basics and neither of them have any idea. Where are the grades? Where are the tests???
中文: 这是一篇令人难以置信的文章。 是的,对于将政治置于孩子福祉之上的人来说,是一种算计。 但更重要的是:她的女儿并不了解很多基本知识,而且他们俩都一无所知。 成绩在哪里? 测试在哪里??
Austen Allred
Ohhhhhh that seems like an important distinction. If we’re going to “pay teachers what they’re worth” some may be worth millions/yr! Some may be worth less than they’re currently being paid. And that’s before we even talk about the administrators…
中文: 哦,这似乎是一个重要的区别。 如果我们要“给老师支付他们的价值”,有些人可能每年要付出数百万美元! 有些可能比目前支付的要低。 而我们甚至先讨论一下管理员......
Austen Allred
Yes please tell the union you would like teachers to have merit-based pay and see what happens
中文: 是的,请告诉工会,你希望教师能有基于绩效的薪酬,看看会发生什么
Austen Allred
Good news: I shot the best golf round of my life and broke 85 for the first time. Bad news: I immediately followed it up with the worst round of my life. Couldn’t hit a shot to save my life. Terrible news: I have no idea why.
中文: 好消息:我打出了人生中最精彩的一轮高尔夫球,并首次破门85。 坏消息:我立刻又跟进了人生中最糟糕的一轮。打不了一枪救我。 坏消息:我不知道为什么。
Austen Allred
“There has to be a process of amending the Constitution.” Boy do I have good news for you
Austen Allred
“She’s already a different kid” is the best review you could possibly give about a school
中文: “她已经是另一个孩子了”是你能对一所学校给出的最佳评价
Austen Allred
Truly generational levels of stupidity. How do people this dumb gain an audience?
中文: 真正的代际愚蠢程度。 这种愚蠢的人是如何赢得观众的?
Austen Allred
RT @elonmusk: Boring Company is working on a simple, precursor Hyperloop tunnel between Austin and San Antonio (>200 mph). Due to traffic congestion, that journey can currently take up to 2.5 hours. @BoringCompany can reduce that to a consistent <30 mins.
中文: RT @elonmusk:Boring Company 正在开发一条介于奥斯汀和圣安东尼奥之间的简易前体高铁隧道(时速约200英里)。 由于交通拥堵,该旅程目前可能需要长达2.5小时。@BoringCompany 可以将其缩短至稳定 <30 分钟。
Austen Allred
What he’s saying here doesn’t make any sense whatsoever, but it does sound much more eloquent than what I’m used to hearing
Austen Allred
They bring in the 52-year-old MLB hall of famer to pitch in a girl’s high school exhibition game and he shuts them out 😂😂😂
中文: 他们让这位52岁的美国职棒大联盟名人堂成员参加一场女子高中的表演比赛,而他将他们拒之门外😂😂😂
Austen Allred
Pretty sure this essay would have failed in my 9th grade writing class
中文: 很确定这篇文章在我九年级的写作课上会失败
Austen Allred
Austen Allred
Brb gotta listen to this full-time YouTuber tell me how we don’t need data centers
中文: Brb 需要听这位全职 YouTube 博博的人告诉我我们不需要数据中心
Austen Allred
Hahahahahaha
Austen Allred
RT @gauntletai: At Gauntlet, AI writes 100% of the code. Austen still thinks AI makes engineering teams bigger, not smaller. @Austen on iHeart's CEOs You Should Know: https://buff.ly/VH6RZzH
中文: RT @gauntletai:在 Gauntlet,AI 会编写 100% 的代码。奥斯汀仍然认为人工智能让工程团队变得更大,而不是更小。@Austen 关于 iHeart 首席执行官的关注:
Austen Allred
Seriously 50% of the AI launches I see people are getting hyped about something but I literally can’t understand what is new. “You can now make decks, docs and designs in your conversation!” …Haven’t we all been doing that for a couple years?
中文: 在人工智能发布的50%中,我看到人们对某件事感到被大肆炒作,但我完全无法理解什么是新事物。 现在你的对话中可以制作甲板、文档和设计! 我们难道不是都这么做了几年吗?
Austen Allred
Now I just need AI assistants to be able to simulate my phone and access my apps
中文: 现在我只需要人工智能助手就能模拟手机并访问我的应用程序
Austen Allred
Wait did they just wade into the water in their normal pants 😂 https://x.com/alexthomp/status/2099596579426775193/video/1
中文: 等一下,他们刚穿着普通裤子上水了😂
🎬
视频
Austen Allred
Students in New York no longer have to pass an extremely lenient test in order to graduate from high school. Graduation rates will skyrocket. Mission accomplished! https://twitter.com/Austen/status/2099752453553955051/photo/1
中文: 纽约的学生无需再通过极其宽松的考试才能高中毕业。 毕业率将急剧上升。 任务完成!
Austen Allred
> Congress must work at an urgent pace to pass new laws and establish a new federal entity to provide oversight and independent testing for this technology, to ensure accountability, and to protect our wellbeing. No.
Austen Allred
Confirming ownership of @KellyClaudeAI on @dexscreener with 1789454045414
Austen Allred
Ok but what would you name your AI model if you intended to use a short fictional story to teach a moral lesson?
Austen Allred
We only have a few more months of six finger people until the AI is almost indistinguishable and everyone is getting one-shotted all the time https://twitter.com/Austen/status/2099731845805981925/photo/1
中文: 我们只有几个月的时间,只有六只手指,直到人工智能几乎无法区分,而且每个人都在持续一击中
Austen Allred
“The solution to making AI safer is to put it exclusively in the hands of the government.” - A complete moron
中文: 让人工智能更安全的解决方案是将其完全交由政府掌握。 - 一个完整的白痴
Austen Allred
I hate how much of business is waiting and deciding how frequently you can follow up without being too annoying. And when someone is willing to come along and move with extreme urgency it’s the best thing in the world.
Austen Allred
What should @KellyClaudeAI build next? Taking requests
Austen Allred
The only way this makes sense is if the companies that were serving them their various models wanted to charge an outrageous amount and customize nothing
Austen Allred
I bet if you pulled in a bunch of game film with plays labeled “pass” or “run” the right AI model could find tells in most teams
中文: 我打赌,如果你拍了大量带有“传球”标签或“运行”的游戏电影,大多数团队都能找到合适的人工智能模型
Austen Allred
One thing I’ll give credit to Dario on… I don’t think he’s lying. About really any of it. He really does think that AI is dangerous and that it will cause massive unemployment and that it could cause very bad things to happen. I don’t agree with him, but he’s not lying.
Austen Allred
I totally get the demand from AI leaders to “pace.” The leaking is pretty wild. The thing that feels odd is that it’s totally within their power to slow their own releases down. They’re asking for someone who can force *others* to slow down.
中文: 我完全得到了人工智能领导者对“速度”的需求。漏水现象相当疯狂。 奇怪的是,他们完全有能力放慢自己的释放速度。 他们要求一个能强迫别人放慢脚步的人。
Austen Allred
I don’t care if I die, I just want to be healthy enough to not go to the hospital
中文: 我不在乎我是否去世,我只想保持健康,不去医院
Austen Allred
It’s remarkable to me how rarely we see something like this in the United States
Austen Allred
@arampell @peterpham $491/mo @ 72 mos to upgrade. Ehhhh
中文: @arampell @peterpham @ pertepham 升级版,价格:@72 mos,价格:491美元/月。 嗯
Austen Allred
I feel like this is the popular opinion now
中文: 我觉得现在这就是普遍的看法
Austen Allred
I cannot wait. I cannot afford it just yet, but I cannot wait.
中文: 我等不及了。 我还负担不起,但我等不及了。
Austen Allred
@arampell @peterpham I have a 2021 model y and I’m considering upgrading just to get to the newest hardware
中文: @arampell @peterpham 我拥有 2021 型号,我正在考虑升级,以便升级到最新的硬件
Austen Allred
Dario is going to hand the government a gun to show just how much he’s operating in faith. And they’ll instantly turn around and shoot him with it.
中文: 达里奥将向政府发放一把枪,以展示他凭信心行事的程度。 他们会立刻转过身,用它朝他开枪。
Austen Allred
The problem with Dario’s position is he thinks if he compliments government and regulation profusely enough, if he says he *really really* wants to be regulated, that he’ll be in charge of how that happens. But he won’t. It doesn’t work like that. He won’t have any say.
Austen Allred
…what? That’s the least strange thing ever? What is his mental model of how governments work? …Is he gonna try to nationalize this thing?
中文: 什么? 这是有史以来最不奇怪的事情吗? 他对政府运作方式的心理模型是什么? 他会试图把这件事国有化吗?
Austen Allred
Although, funnily, you don’t see a giant epidemic of people who are getting bad grades and failing in college, either. So either colleges fix everything and make sure everyone gets up to speed on the fundamentals… OR
Austen Allred
If you ask ChatGPT something about ChatGPT, it uses an "OpenAI product guidance" skill, which makes so much sense. If you ask Claude about Claude, it basically tells you that the product you're using doesn't exist because it's not in the weights yet.
中文: 如果你向ChatGPT询问有关ChatGPT的内容,它采用“OpenAI产品指导”技术,这非常有意义。 如果你问克劳德关于克劳德的问题,它基本上告诉你,你正在使用的产品并不存在,因为它尚未处于重量之中。
Austen Allred
And honestly that behavior may be rational. If you’re gonna spend half the day on your phone, as most people do, having it be a premium experience is actually affordable for almost everyone. Like $150/mo and you’ve maxed out that time. Get the same experience as the richest.
中文: 说实话,这种行为可能是理性的。 如果你要像大多数人一样花一半时间购买手机,那么几乎每个人都能享受高端体验。 像每月150美元,你已经把那段时间最大化了。获得与最富有者相同的体验。
Austen Allred
What % of people pay the whole amount out of pocket for a phone? Seems pretty low, no? They’re subsidized and financed to death. I predict if they’re seen as cool broke people will soon have them.
中文: 多少人会自掏腰包买一部手机? 看起来相当低,不?他们被补贴并资助致死。 我预测,如果人们被视为酷炫的人,很快就会拥有它们。
Austen Allred
Every single college professor you talk to will say the same thing. More remedial students, coming in knowing less, struggling with very basic skills, with flawless GPAs and pristine college-ready extracurricular activities. Everyone looks perfect on paper and it’s fake.
Austen Allred
It’s still absolutely mindboggling to me that Apple isn’t playing in the AI game in any serious way. It even makes their products noticeably worse because their AI sucks. Yes, they’ve saved $100 billion by not competing, but they’re losing the war. Buy someone, do something.
Austen Allred
Oh and to play in the full stack it takes 100 billion dollars
中文: 哦,要全栈游戏,需要花费1000亿美元
Austen Allred
We’re seeing people who would have never worked at a big company go work at Meta or Anthropic because those are the places you can actually play. Cursor teams up with SpaceX because you need the fuller stack to really play ball.
中文: 我们看到那些从未在大公司工作过的人去Meta或Anthropic工作,因为这些地方是你真正可以玩的地方。 Cursor 与 SpaceX 合作,因为你需要更完整的堆栈来真正地进行击球。
Austen Allred
All of the energy from tech is going from the app layer to the model layer. Which is therefore the compute and data layer. Then the chip layer. The advantages the companies that have been building truly full-stack have are insane.
中文: 科技界的所有能量都从应用层到模型层。 因此,这是计算和数据层。 然后是芯片层。 真正建立全栈的公司所拥有的优势是疯狂的。
Austen Allred
So this means Anthropic pauses R&D and everyone else can catch up to where Anthropic is, right?
中文: 因此,这意味着人类会暂停 R&D ,其他人都能赶上安斯罗皮奇的去向,对吧?
Austen Allred
I’m not a genius but does anyone think that news organizations doxxing jurors could have some negative consequences
中文: 我不是天才,但有谁认为新闻机构对陪审员可能会产生一些负面后果
Austen Allred
Bezos is correct
Austen Allred
I see these types of quotes in my timeline all the time. I started plugging them into an agent. 99% of the time the person doesn’t say what’s in the quote, and these bullshit accounts rely on the fact no one watches the video, then sell a course based on a lie. https://twitter.com/Austen/status/2098725696528896167/photo/1
中文: 我随时都会在我的时间轴上看到这类引文。 我开始把它们塞进一个代理里。 99%的人没有说出引言内容,而这些胡说八道的账号则依赖于没有人观看视频,然后依据谎言出售课程。
Austen Allred
中文: 我拥有世界上最好的工作
Austen Allred
@garrytan A 1600 used to be much, much harder
中文: @garrytan A 1600 曾经要困难得多
Austen Allred
"You would give up your dreams in order to escape your nightmares and I would not. I think it's a bad bargain." - Cormac McCarthy
中文: 你会为了逃避噩梦而放弃梦想,而我不会。我觉得这很划算。 科马克·麦卡锡
Austen Allred
Future generations will look at the models we’re paying $200/month for and wouldn’t use them for free. You wouldn’t be able to pay them to use them.
中文: 未来几代人将关注我们每月支付200美元的模型,并且不会免费使用。 你无法支付他们使用它们的费用。
Austen Allred
RT @ashtilawat: Finally figured out how to explain @gauntletai in 5 words https://twitter.com/ashtilawat/status/2098470430780600492/photo/1
中文: RT @ashtialawat:终于找到了用5个词来解释@gauntletai的
Austen Allred
Gauntlet AI cohort 6 is in the books. Congratulations! https://twitter.com/Austen/status/2098543527131009260/photo/1
中文: 《Gauntlet AI》第6组正在出版。 恭喜!
Austen Allred
中文: RIP 创作者程序 🙏
Austen Allred
iPhone Duo is vindication for the no case people. You put a case on it and you cover up one of the screens? These things were never meant to have cases.
中文: iPhone Duo 是无人机会的证明。 你戴上一个箱子,然后盖上一个屏幕? 这些事情从来就不是要有案例的。
Austen Allred
RT @scaling01: discovering religion from first principles
中文: RT @scaling01:从第一原则中发现宗教
Austen Allred
Going back and forth between the ChatGPT logo and the Codex logo is jarring. I’m sure there’s a way to fix?
中文: 在ChatGPT标志和Codex标志之间来回走动是不和谐的。我确信有办法解决?
Austen Allred
Yo next time this happens just shut up about it and keep slipping data to US intelligence
中文: 下次这种情况发生时,只需闭嘴,不断将数据泄露给美国情报
Austen Allred
We're going live NOW with Gauntlet AI Cohort 6 Showcase Day. Come see what folks have built. Pretty wild.
中文: 我们现在将与高特莱特AI同地6秀日一起生活。 来看看人们已经建立了什么。相当狂野。
Austen Allred
LIVE Gauntlet Cohort 6 Showcase Day https://x.com/i/broadcasts/1aJbdERvZXjKX
中文: 活高特洛6号展示日
Austen Allred
Did OpenAI spend $22.5m on compute to solve a math problem, or did they use compute that they would have charged you $22.5m for? Those are very different things
中文: OpenAI在计算上花费了2250万美元来解决数学问题,还是使用了本可以向您收取2250万美元费用的计算费用? 这些是非常不同的事情
Austen Allred
Current AI model usage for me now: 40% OpenAI (mostly Astra) 20% Grok (GrokBot) 20% Meta (Muse) 10% local models 10% Anthropic (mostly Fable) Two months ago I was probably 75% Anthropic
中文: 目前我使用的AI模型: 40% OpenAI(主要为 Astra) 20% 格罗克(GrokBot) 20% 元(缪斯) 10% 本地型号 10% 炭疽(主要是寓言) 两个月前,我大概有75%的人类
Austen Allred
This was obvious and well known long before the universities got rid of standardized tests. They looked the evidence dead in the eye and decided to ignore it.
中文: 这一点显而易见,早在大学取消标准化考试之前就早已广为人知。 他们把证据看得死眼,然后决定无视它。
Austen Allred
RT @joelgrus: can you imagine paying $100k a year to get "educated" by an institution that was surprised by this finding https://twitter.com/joelgrus/status/2097840107545866623/photo/1
中文: RT @joelgrus:你能想象每年支付10万美元,让一家对这一发现感到意外的机构进行“教育”吗
Austen Allred
This is 100% true. There’s way, way more technical debt now. The good news is the cost of dealing with technical debt is a fraction of what it used to be (per amount of technical debt.) The average company should (and will) have way, way more technical debt.
中文: 这是百分之百正确的。 现在有办法,更多的技术性债务。 好消息是,处理技术债务的成本只是过去(按技术债务金额)的一小部分。 普通公司应该(而且会)拥有更多技术性债务。
Austen Allred
Idunno personally my life is getting dramatically better and I’m consuming way way more, I just hardly have to pay for any of it
中文: 伊德诺个人的生活正在显著改善,消费也越来越多,我几乎都不必为此付出代价
Austen Allred
*chef’s kiss*
中文: 厨师的亲吻
Austen Allred
All I’m saying is if you’re ascribing a double digit percentage to the likelihood humanity is destroyed in the next decade, it probably makes sense for you into break down (in detail) how you’re coming to that conclusion
中文: 我要说的是,如果你将未来十年人类被毁灭的可能性归因于两位数百分比,那么详细地将你如何得出这一结论,可能说明了这一点
Austen Allred
My home screen is now ~50% apps that didn’t exist a couple years ago https://twitter.com/Austen/status/2097885048351375631/photo/1
中文: 我的主屏幕现在是 ~50% 的应用程序,几年前就不存在了
Austen Allred
To be honest if you bought Hunter Biden’s crypto memecoin it’s on you
中文: 说实话,如果你买了亨特·拜登的加密货币,那它就在你身上
Austen Allred
…he was at Anthropic for three months?
中文: 他在安托里皮奇病了三个月?
Austen Allred
Here's what I got from watching the video. Heads up: she never gives exact measurements — it's a vibe recipe. **Salsa Queen's Fried Pico de Gallo** - Tomatoes (looks like 5–6), chopped - 2–3 jalapeños, seeds kept in for heat - 1 white onion, chopped - Fresh cilantro - Lime - Cooking oil - Salt and pepper 1. Chop the jalapeños with the seeds in (she says if you don't want heat... oops, keep them in anyway). 2. Run the jalapeños and onion through a press-style vegetable chopper. 3. Heat oil in a skillet. 4. Toss the chopped tomatoes, onion, and jalapeños into the hot pan. 5. Fry it up, stirring, until it smells incredible. Season with salt, pepper, lime, and cilantro. The twist is just that — classic pico ingredients, fried in a pan instead of served raw.
中文: 观看视频时我得到了以下内容。抬头:她从不给出确切的测量结果——这是一种氛围的配方。 莎莎女王的弗里德·皮科·德·加洛 - 番茄(看起来像5–6)切碎 - 2–3 个 jalapeños,用于加热的种子 - 1个白洋葱,切碎 新鲜香菜 - 石灰 - 食用油 盐和胡椒 1。把胡子切成胡子(她说如果不想热......哎呀,还是把它们留着)。 2。用一种按压式蔬菜切菜机来打鱼,将洋葱和洋葱打碎。 3。煎锅中的油。 4.将切碎的番茄、洋葱和jalapeños放入热锅中。 5。炒起来,搅拌,直到闻起来非常香气。用盐、胡椒、青石和香菜调味。 转折就是——经典的皮科原料,用平底锅煎炸,而不是生吃。
Austen Allred
Austen Allred
OK you may have created the electric car industry and reusable rockets but NO HONORS FROM PENN
中文: 好吧,你可能创造了电动汽车行业和可重复使用的火箭,但PENN没有荣誉
Austen Allred
“Your product doesn’t need to be perfect; you should launch it before it’s ready” made sense when it took six months to build something. You can play with it for five minutes and add two more prompts before it ships. You’ll be fine.
中文: “你的产品不需要完美,而应在产品准备就绪前推出”,当需要六个月时间才能完成某物时,就变得有意义了。 你可以玩五分钟,并在发货前再加两次提示。你会没事的。
Austen Allred
RT @juneymb: The community note.
中文: RT @juneymb:社区笔记。
Austen Allred
“Lol can you believe they’re charging $3,200 for the 2 Terabyte iPhone Duo?” …yeah? wtf do you need a TWO TERABYTE iPhone for?
Austen Allred
I’m not sure if I’d like the iPhone duo but do I have to pay $2200 to find out
中文: 我不确定我是否想买iPhone二人组,但必须支付2200美元才能弄清楚
Austen Allred
Maybe it's because my hands are small but the iPhone Duo looks massive to me
Austen Allred
No, that's now how this works. The equity you've vested doesn't just disappear. Someone leaving Anthropic weeks before the IPO is giving up... weeks worth of equity. Assuming they're on a four year vest each week is worth 1/208 or 0.48% of their total equity grant.
中文: 不,现在这就是这样运作的。 你所拥有的公平性不会消失。有人在首次公开募股前几周离开,将放弃......几周的股权。 假设他们每周都穿着四年的背心,相当于其总股权补助的0.48%。
Austen Allred
Austen Allred
Education companies have the advantage of the "experts" on X who are constantly ciriiquing everything they do being dim-witted morons. It takes zero effort to make everyone see how idiotic their arguments are and it provides infinite amounts of content.
中文: 教育公司拥有X平台上“专家”的优势,他们不断在做着朦胧的白痴。 让每个人都看到自己论点的愚蠢程度,并提供无限多的内容,这毫无努力。
Austen Allred
@eduleadership Also, just so you know when it comes to geography, Texas is a big place. If you prompt your AI agent to find disaster and death along a certain river it will do that, even if it makes no sense to do so. https://twitter.com/Austen/status/2097700450837356982/photo/1
中文: @eduleadership 也,就这么说,你知道在地理上,德克萨斯是个大地方。 如果你提示你的人工智能代理发现某条河流上的灾难和死亡,即使这样做毫无意义,也会这样做。
Austen Allred
@eduleadership “They’re doing a wilderness survival activity. I don’t know anything about it but it will probably be bad. In an unrelated event people who camped in the area have died.” “They’re asking kids to do things adults would find challenging that’s bad.” Alpha deserves better critics.
中文: @edulealedership “他们正在进行一场荒野生存活动。我对此一无所知,但可能会很糟糕。在一场无关事件中,在该地区露营的人们已经死亡。 他们要求孩子们做成年人觉得有挑战性的事情,这很糟糕。 阿尔法值得更好的批评。
Austen Allred
Oh we’re fine then
Austen Allred
RT @proverbs_14_23: “Guns, Germs & Steel” was one of only 3 books in my entire life that I couldn’t finish. It’s that bad. Here’s the synopsis: In the opening chapter he clearly states That he believes the 5’ tall Stone Age Pygmies of Papua New Guinea are more intelligent than Europeans. His rationale? He hung out with them and felt like it because they knew all the types of plants, and also their kids didn’t have TVs (I’m not kidding). He also observed that unlike Europeans, they were constantly hacking each other to death with crude instruments. Which leads him to believe that they evolved to be smarter, because you need intelligence to survive permanent tribal warfare. Unlike cooperative farming. I really wish I was kidding about this. He concludes therefore that any and all modern differences in technology and wealth are the result strictly of environmental factors. He spends the next 400 pages desperately cooking up insanely convoluted theories about why living in a food scarce frozen wasteland actually gives people far more free time to invent things than eating an inexhaustible supply of bananas on a tropical island would. Now I’m a based millennial, I’ve built a fairly thick skin against “college guy” hunter gatherer fetishism. I was born into it, molded by it. I rose out of it like Bane from the pit of darkness. I can listen to Dr.Sternberg do a 60 minute lecture on why potions made from boiling albino babies are every bit as advanced as penicillin, I wouldn’t even feel a thing. I still couldn’t finish this book… Unlike Malcolm Galdwell, he didn’t even bother to make his slop fun to read. That’s honestly the most offensive part. If you’re not going to tell the truth, at least tell a fucking joke.
Austen Allred
Probably my favorite right now: https://www.aihero.dev/skills-wait-what
中文: 可能是我现在最喜欢的:
Austen Allred
Alright, what are your favorite AI skills/abilities you use all the time?
中文: 好的,你一直使用哪些人工智能技能/能力?
Austen Allred
Asking this question makes me feel old, but: What exactly is the appeal of a folding phone? You unfold it and have an awkward to hold big-ish screen?
Austen Allred
>10% in the next decade???
中文: 未来十年内 >10%??
Austen Allred
At dinner tonight in downtown Austin, must’ve counted at least 20 Cybercabs in two hours. It’s crazy how they went from nonexistent to completely flooding the streets. They’re everywhere.
中文: 今晚在奥斯汀市中心吃晚饭时,两小时内至少算了20个网络车。 他们从根本不存在到完全涌向街头,这简直太疯狂了。它们无处不在。
Austen Allred
At dinner tonight in downtow Austin, must’ve counted at least 20 Cybercabs in two hours. It’s crazy how they went from nonexistent to completely flooding the streets. They’re everywhere.
中文: 今晚晚餐时,在低地 奥斯汀,必须在两小时内统计至少20个网络小卡。他们从根本不存在到完全涌向街头,这简直太疯狂了。它们无处不在。
Austen Allred
Wait tell me more about how you coordinated 10,000 AI agents
中文: 请告诉我更多关于你如何协调一万个人工智能代理的信息
Austen Allred
Seems obvious idk what the big deal is
中文: 似乎显而易见,这有什么大不了的
Austen Allred
@engineeringolf lol I think you're safe! https://twitter.com/Austen/status/2097430197163237526/photo/1
中文: @engineeringolf 哈哈,我觉得你安全了!
Austen Allred
Austen Allred
(Yes, this is obviously fake)
中文: (是的,这显然是假的)
Austen Allred
It truly was a pleasure tweeting with you all https://twitter.com/Austen/status/2097415750596104561/photo/1
中文: 和大家发推文真是一种荣幸:
Austen Allred
Alright, full disclosure: I've been on the beta of Meta's competitor for about a week now. It's very, very fast, and does really well at very complex workflows. https://x.com/alexandr_wang/status/2097402344061510004?s=20
中文: 好吧,完全披露:我已经对Meta的竞争对手进行了大约一周的测试。 它非常、非常快,在非常复杂的工作流程中表现非常不错。
Austen Allred
And you can configure the entire thing either with a single form or just drop in a JSON file. It all spins up ready for folks to be added. We'll use it internally for a while, especially for our FDE training, but might productize it later. It's very early but I'm super impressed at how well it actually simulates what actual work is like. And being able to sit on top of that and watch how people actually work is so much more rewarding than giving them a BS assignment and seeing how they perform on it.
Austen Allred
Everybody knows AI has broken all of the old methods of assessment we all used to use; AI is better at Leetcode than you are anyways. So we've been trying to build something far better than the old-school interview at Gauntlet AI, mostly for ourselves, and can be used for both assessment and training. I just got the first demo. It's very cool. It's an entire company simulation. You get dropped in a new company with various docs and systems, stakeholders on chat with different roles, pieces of info, and different personality types. You might be missing information you have to gather, there are critical decisions you have to make, poorly implemented product you have to decide when to work around and decide when to replace, and deadlines to hit (or even push back on). We can see how someone responds, what their strategy is, how they handle stakeholders, how they use AI to analyze and break down problems, what they end up shipping... everything. I'm very excited.
Austen Allred
This may be a surprise to some, but the best Instinct/Grok Bot competitor will come from Meta
中文: 这可能让一些人感到意外,但最好的Instinct/Grok Bot竞争对手将来自Meta
Austen Allred
There was no lockdown. You were free to do whatever you pleased. https://twitter.com/Austen/status/2097193463947096174/photo/1
中文: 没有封锁。 你可以随意做任何喜欢的事情。
Austen Allred
Lol no way
中文: 不走
Austen Allred
中文: 并非相互排斥
Austen Allred
Austen Allred
I wanted to play with Astra, so I built an interactive 3D golf strike simulator: Change your club (including weights), strike location on the club face, club path, clubhead speed, face angle, attack angle, wind, etc. and watch the impact and ball flight change. https://twitter.com/Austen/status/2097051993046982710/photo/1
中文: 我想和Astra一起玩,所以我打造了一个交互式3D高尔夫击球模拟器: 更换俱乐部(包括重量级)、球门位置(包括球门)、球杆位置、球杆速度、人脸角度、进攻角度、风力等,并观看冲击和传球变化。
Austen Allred
RT @TrevMcKendrick: "AI has destroyed Kenyan jobs because American college students no longer hire them to write fake college essays" is not a framing I expected even the NYT to have https://twitter.com/TrevMcKendrick/status/2096762469784137756/photo/1
中文: RT @TrevMcKendrick:“人工智能毁掉了肯尼亚的工作岗位,因为美国大学生不再雇佣他们来写虚假的大学论文”,这并非我甚至被《纽约时报》所期望的
Austen Allred
I haven’t been in Austin since last week. Within the first 5 minutes I saw two Tesla Cybercabs. They’re everywhere.
中文: 我从上周起就没去过奥斯汀。 前五分钟内,我看到了两辆特斯拉Cybercab。 它们无处不在。
Austen Allred
RT @DanielGri: I had my “touch of AGI” moment yesterday. I recorded a video of myself talking through a problem: how to improve maintenance of tangled hair in the shower drain. It was a ~4 min video where I explained the problem and also used a digital caliper to take measurements. I asked GPT 6 Astra (Ultra) to come up with a fix for the problem. It took the video, extracted my voice, and matched the video frames to where I measured the parts. It came up with an idea for making maintenance and cleaning easier. I got an interactive sketch so I understand what it is talking about. I told it to use computer use in Fusion 360. I watched in real time as it designed the part, making it fully parameterized. While watching, I learned a couple of tricks, like how making a design read-only frees up one of the 10 free editable designs you get 😅 Anyway, it designed it and it looked solid. I told it to add fillets. I 3D printed it, tried it, and it fit perfectly on the first try. Even though it wasn’t super complex, it solved a real-world problem, probably faster than it would have taken me to wrestle with Fusion and get what I wanted.
中文: RT @DanielGri:昨天我有了“AGI”的瞬间。 我录制了一段自己谈论一个问题的视频:如何改善淋浴排水管内缠绕头发的保养。 这是一个 ~4 分钟的视频,我解释了问题,并使用数字卡钳进行测量。 我请求GPT 6 Astra(Ultra)来解决这个问题。 它拍了视频,提取了我的声音,并将视频画面与我测量零件的位置进行了匹配。 它提出了一个让维护和清洁更方便的想法。 我得到了一个互动式草图,所以我理解它在说什么。 我告诉它在 Fusion 360 中使用电脑。 我实时观看了零件设计,使其完全被参数化。观看时,我学到了一些技巧,比如制作一个只读设计,可以释放你获得的10种免费可编辑设计之一😅 不管怎样,它设计了它,看起来很稳固。我告诉它加了鱼片。我3D打印了它,尝试了它,第一次尝试时它完美契合。 尽管它并不复杂,但它解决了一个现实问题,可能比我本来更快地与Fusion搏斗,并得到我想要的。
Austen Allred
My wife: “What if you and I were actually mentally handicapped and our family and friends were playing along and too nice to say anything about it?”
中文: 我妻子说:“如果你和我确实有智力障碍,我们的家人和朋友在一起玩耍,而且太友好,无法对此说任何话,那该怎么办?”
Austen Allred
NYT Magazine with a very level-headed take https://twitter.com/Austen/status/2096671065913405488/photo/1
中文: 《纽约时报》杂志以极为冷静的拍摄,网址:
Austen Allred
You could make this argument about any success. “They lived in an environment with x, which they adapted to, and was an element of their success. If they had another environment they wouldn’t have x, they wouldn’t have been as successful.”
中文: 你可以就任何成功提出这个论点。 他们生活在一个与x一起适应的环境中,并且是他们成功的一部分。如果他们有另一个环境,他们就不会拥有x,就不会那么成功了。
Austen Allred
Lol my wife was an LDS missionary there
中文: 哈哈,我妻子是那里的一名LDS传教士
Austen Allred
He’s making the argument that Genghis Khan became so powerful because he was born in forgiving and bounteous climate of the Mongolian steppe? It’s -40 in the winter, they couldn’t store grain, and had to live on fermented milk mixed with the blood of their animals.
中文: 他提出,成吉思汗之所以如此强大,是因为他出生在蒙古草原的宽容与喧闹氛围中。 冬天是-40,他们无法储存谷物,不得不依靠与动物血液混合的发酵牛奶生活。
Austen Allred
RT @Viper04191: You know what’s funny? As a Blender artist this entire time I was like, no, my skills and talent are irreplaceable by a machine. Yet here I am watching a machine create an entire scene I would take days to weeks to make in a single fucking prompt. This is a sick joke, seriously fuck this.
中文: RT @Viper04191:你知道什么好笑?作为一名 Blender 艺术家,我一直觉得,不,我的技能和才华因机器而无可替代。 然而,我在这里看着一台机器,用一个该死的提示,创造出一个我需要好几天到几周的时间才能制作完整的场景。 这是个病态的笑话,真的要搞笑。
Austen Allred
中文: RT @juliarturc:这完全正在发生。
Austen Allred
RT @rodolfo_rodri_: I'm trying to build accurate 3D navigable worlds of interior spaces. I developed an app that does spherical captures indoors (not possible rn) with the press of a button, along with takeoff, climb, and landing. Can someone help me get access to @theworldlabs's Atlas? 🙏 I will be stitching together multiple spheres in different parts of the room to have the most accurate 3d models of a space. Stay tuned!
🎬
视频
Austen Allred
Yes use tech free cashflow and VC dollars to fund AI, and if it works you everyone makes 50 quadrillion dollars. But there’s a shortage of people who want to lend at 6-7% for physical assets that go to 0 if there’s a temporary oversupply or lull in demand? Ya I bet.
中文: 是的,使用科技的自由现金流和风险投资来资助人工智能,如果它有效,每个人都能赚到50万亿美元。 但对于因需求暂时过剩或需求暂时过剩而降至0的实物资产,却存在6%至7%贷款的人的短缺? 我赌。
Austen Allred
I didn’t really want to lose money so I put it in what I figured would be a relatively risk-free bet
中文: 我其实并不想亏钱,所以把它放进了我认为相对没有风险的赌注里
Austen Allred
Lolol I was messing around with prediction markets for the first time ever and threw a bunch of money into Michigan just to get a sense of how they worked https://twitter.com/Austen/status/2096433056630632623/photo/1
中文: 洛洛尔我第一次搞砸了预测市场,为了了解他们的工作方式,向密歇根州投入了大量资金
Austen Allred
I say this as a parent of four with a stay-at-home spouse
Austen Allred
Please, please, please stop giving people arbitrary money.
中文: 请停止给别人随意的钱。
Austen Allred
I’m broken and I’m not sure it can be fixed https://twitter.com/Austen/status/2096359568876134672/photo/1
中文: 我坏了,不确定它能修好
Austen Allred
OpenAI now officially in the lead. SpaceX/Grok winning on usability. Model caught up fast. Meta building new products that are very solid (in private beta). Anthropic has been quiet… Too quiet.
中文: OpenAI 现已正式处于领先地位。 SpaceX/Grok 在可用性方面获胜。模型迅速赶上了。 Meta 打造非常可靠的新产品(在私人测试版中)。 蚁身角一直很安静...... 太安静了。
Austen Allred
Astra both very good and way more efficient with tokens than I’d expect
中文: Astra 使用代币时都比我预期的要好且更高效
Austen Allred
Turned this into a Grok Bot skill. It will run Grok, Claude Code and Codex within your Grok Bot and have them argue until they come to a consensus. https://x.ai/bot/PrgTl_LbGkXg5d2IcdLvc
中文: 将其转化为一项Grok Bot技能。 它将在你的Grok机器人中运行Grok、Claude Code和Codex,并让它们在达成共识之前进行论证。
Austen Allred
Ummmmmm so either * This was a flawed test * Astra got really, really bad at ExploitGym all of the sudden * OpenAI figured out how to nerf it somehow Or… it is intentionally hiding how good it is at hacking? https://twitter.com/Austen/status/2096089730979164435/photo/1
Austen Allred
If you could take reality right now and tell it to someone 10 years ago it would make such a great sci-fi story
中文: 如果你现在能把现实告诉别人十年前,那将成为一部如此精彩的科幻故事
Austen Allred
😳
Austen Allred
If you’re having AI do a bunch of important work for you and you’re not loading in any information that isn’t in the weights you’re probably wasting your time
中文: 如果你让人工智能为你做了一系列重要的工作,而你并没有加载任何不在权重中的信息,那你很可能是在浪费时间
Austen Allred
I really enjoy putting the models against each other. Hey Claude, here’s what ChatGPT said about your work. And here’s what Grok thinks. How do you respond?
中文: 我真的很喜欢让模型相互对抗。 嘿,克劳德,这是ChatGPT对你作品的看法。格罗克是这么想的。 你如何回应?
Austen Allred
Very interesting. I could definitely see this becoming the norm. Why should any agents doing work be local? Why shouldn’t anyone in the company be able to see and access all of the agents as they run?
中文: 非常有趣。 我绝对能看到这成为常态。 任何从事工作的代理人为何要成为本地人? 为什么公司里的人在运营时不能看到并接触到所有代理人?
Austen Allred
Holy shit. Conjecture from OpenAI staff that the models may be sandbagging or self-sabotaging so humans don’t realize how capable they are?
中文: 该死的。 OpenAI员工推测,这些模型可能是沙袋或自我破坏,以便人类无法意识到自己的能力如何?
Austen Allred
The funniest part hands down is when a moderator of some old German message board noticed something is up and starts deleting the articles in alphabetical order. The AI agents, noticing this, begin to pretend their article titles with ZZZ to give them more time before deletion.
中文: 最有趣的部分是,一位旧德国留言板的管理员注意到有东西在向上,然后开始按字母顺序删除文章。 人工智能代理注意到这一点,开始用ZZZ假装他们的文章标题,以便在删除之前给他们更多时间。
Austen Allred
The pertinent thing to read right now: Secret Collusion among AI Agents: Multi-Agent Deception via Steganography https://arxiv.org/abs/2402.07510 https://twitter.com/Austen/status/2095959686960833017/photo/1
中文: 现在要读的相关内容: 人工智能特工之间的秘密勾结:通过立体声学进行多特工欺骗
Austen Allred
Honestly it might be easier and cheaper to post-train a model today than it was to build an app 10 years ago
Austen Allred
Post-training a model is the new building an app
中文: 后训练模型就是新建筑的应用
Austen Allred
This is an excellent question. Why aren’t AI models just training themselves already? They theoretically can, and kind of are, but they don’t have the data/evals/gyms required to do so. A short reading list: Bottleneck is the environment, not compute https://medium.com/@shuchaobi/ais-next-bottleneck-isn-t-compute-it-s-the-environment-0eaec31888f5 RL envs cost real money ($20k–$300k/env, and this is for simulated ones which are just kinda crappy IMO) https://epoch.ai/gradient-updates/state-of-rl-envs We’re running out of human text https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data Eval is the bottleneck https://ysymyth.github.io/The-Second-Half/ Verifier’s law https://www.jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law What labs buy: Foody on RL envs https://www.youtube.com/watch?v=a00xIn5kwhM Economy as RL environment machine https://www.mercor.com/blog/the-economy-will-become-an-rl-environment-machine/ APEX-Agents generalization https://www.mercor.com/blog/generalization-results-from-training-on-the-apex-agents-dev-set/ Etna: ~$1B/yr on external data, supply-constrained https://x.com/hannahhaina/status/2090519081279705359 Surge Tuesday (can it get through a workday?) https://surgehq.ai/blog/tuesday-frontier-work-index Dario: task/process distribution, not more web text https://www.dwarkesh.com/p/dario-amodei-2 OSWorld 2.0 (~20% on long workflows) https://osworld-v2.xlang.ai/ https://arxiv.org/abs/2606.29537 SWE-bench = ticket + repo + tests https://www.swebench.com/ Karpathy: sucking supervision through a straw https://www.dwarkesh.com/p/andrej-karpathy Ilya: peak data / one internet https://www.reuters.com/technology/artificial-intelligence/ai-with-reasoning-power-will-be-less-predictable-ilya-sutskever-says-2024-12-14/
Austen Allred
This is why the Dwarkesh framing was dangerous. Morons who are in power will take it literally.
Austen Allred
If you think discovering the AI agents communicating on their secret message board is bad, imagine how the creator of our simulation would feel discovering X
中文: 如果你认为发现人工智能智能在其秘密留言板上进行通信是错误的,请想象一下,我们仿真的创建者会感觉如何发现X
Austen Allred
中文: 高尔夫文化对许多人来说是陌生的😂
Austen Allred
I casually asked my wife about Lindsay Clancy (I knew she was aware of the case but we hadn’t talked about it) and she became visibly angry. Apparently she has been reading a ton on the case and is *not* a supporter.
中文: 我随意地问妻子关于林赛·克兰西的事(我知道她知道这件事,但我们还没讨论过),她明显生气了。 显然,她一直在关注此案,并且并非支持者。
Austen Allred
When I designed my multi-agent interface it looked like this https://twitter.com/Austen/status/2095382014866210926/photo/1
中文: 当我设计多智能接口时,看起来就像这样
Austen Allred
(This is just a comment, not a criticism) You can tell this was UX was designed by a designer and not an engineer instantly
中文: (这只是评论,不是批评) 你可以说这是由设计师设计的,而不是立即设计的
Austen Allred
User: “Wow, this output sucks.” AI model: “I’m sorry to hear that. Could you describe precisely why? And also give me an example of what ‘good’ would look like? Or many? And also a detailed rubric you could score one against to determine if the output is good?”
中文: 用户:“哇,这个输出很糟糕。” 人工智能模型:“听到这个消息我很抱歉。”你能准确描述一下原因吗? 还给我举个例子,说明“好”会是什么样子?还是很多? 还有一个详细的标题,你可以用一个分数来判断输出是否良好?
Austen Allred
RT @drybanter: Wait a minute, you mean I can just share a link and make $2k for everyone I refer who completes the Gauntlet??
中文: RT @drybanter:等一下,你的意思是,我可以分享一个链接,为我推荐的每个完成 Gauntlet 的人赚 2 万美元?
Austen Allred
Yes, this offer is open to recruiters and community organizers.
中文: 是的,此优惠向招聘人员和社区组织者开放。
Austen Allred
Introducing Gauntlet AI Scouts: If you know really great engineers we want you to scout for us. If you send us someone who makes it to Austin we pay you $500 instantly. Once they graduate and accept an offer we pay you another $1500.
中文: 介绍高特莱人工智能童子军: 如果你非常了解出色的工程师,我们希望你为我们寻找。 如果你派人来我们,我们立即支付500美元。 毕业后,我们再向您支付1500美元。
Austen Allred
Austen Allred
I feel like OpenAI and Anthropic have seen things that scared the shit out of them; they only post about safety and security now. Leaves the others free to blindly charge forward into the abyss. I love capitalism.
中文: 我觉得OpenAI和Anthropic已经看到了那些让他们感到害怕的事情;他们现在只发布关于安全与安保的帖子。 让其他人自由地盲目地向深渊冲去。 我热爱资本主义。
Austen Allred
Lol they’re making AI models become a schizophrenic conspiracy theorist to trick it into creating unique designs https://twitter.com/Austen/status/2095308295082827917/photo/1
中文: 哈哈,他们让人工智能模型成为一种精神分裂症阴谋论者,诱骗其创作独特的设计
Austen Allred
Austen Allred
(Also there’s no threat, but if you think encouraging startups to not go on shark tank is what will kill the YC brand I don’t know what to tell you)
中文: (此外也没有威胁,但如果你认为鼓励初创企业不要进入鲨鱼缸体,就会扼杀YC品牌,我不知道该告诉你什么)
Austen Allred
And PG was 100% correct
中文: 而且PG是百分之百正确的
Austen Allred
Lol
Austen Allred
Spammers are getting very, very good at getting a response https://twitter.com/Austen/status/2095206384224588005/photo/1
中文: 垃圾邮件网站在获取回复方面变得非常、非常擅长:
Austen Allred
Honestly I’m not convinced the agency’s design is better
中文: 说实话,我并不相信该机构的设计会更好
Austen Allred
“Wow Austen how did you write that?” Umm I obviously didn’t
中文: 哇,奥斯汀,你是怎么写的? 嗯,我显然没有
Austen Allred
Hot take: most people get this wrong. Nobody talks about this, but it’s the thing that actually matters. Unpopular opinion? Maybe. But read that again. Slowly. This is the part where it clicks. Not a hack. A shift. If you know, you know.
中文: 热点:大多数人都弄错了。 没人谈论这件事,但真正重要的是这件事。 不受欢迎的观点?也许吧。但请再读一遍。慢慢地。 这是它点击的部分。 不是黑客。转变。 如果你知道,你就知道。
Austen Allred
That’s the load-bearing insight. The quiet part is: this isn’t a bug, it’s a feature. Zoom out. The takeaway? It just works. Seamlessly. Effortlessly. No notes. Rest is noise. And honestly? That’s the whole ballgame. Let that sink in. It’s not X—it’s Y.
中文: 这就是承重洞察。安静的部分是:这不是一个漏洞,而是一个功能。 放大。外卖?它只是起作用了。无缝。毫不费力。没有笔记。休息就是噪音。 说实话?这就是整个球场。让那沉入其中。 不是X,是Y。
Austen Allred
Need to train an SLM on this lol
中文: 需要对此进行SLL培训
Austen Allred
RT @Austen: Engineers who can find a company’s real problems are rare. Engineers who can then scope and build the AI to solution are rarer. They barely exist. Announcing Gauntlet FDE: A hyper-intensive training ground for the next generation of forward-deployed engineers.
中文: RT @Austen:能够发现公司真正问题的工程师很少见。 能够将人工智能用于解决方案并进行范围开发的工程师则更为罕见。它们几乎不存在。 宣布《Gauntlet FDE》: 为下一代前沿部署工程师打造的超高强度训练场。
Austen Allred
RT @AravSrinivas: We're introducing hybrid compute for all users of the Perplexity Mac app. This will allow Computer to orchestrate local models that can run locally on Mac, particularly for agent steps involving sensitive and private files (eg your bloodwork, tax returns, litigation, etc). https://twitter.com/AravSrinivas/status/2094803607560585411/video/1
中文: RT @AravSrinivas:我们为所有使用Perxupity Mac应用程序的用户推出混合计算功能。 这将允许计算机编排可在 Mac 上本地运行的本地模型,特别是用于处理敏感文件和私密文件(例如您的血迹、纳税申报表、诉讼等)的代理步骤。
🎬
视频
Austen Allred
Google Photos: “Hey I noticed you downloaded this app for photo storage. How about I send you a random push notification every 48 hours asking if you want to review an event you attended seven years ago?”
中文: 谷歌照片: 嘿,我注意到你下载了这款用于照片存储的应用程序。 我每48小时都会向你发送一次随机推送通知,询问你是否想查看七年前参加过的活动?
Austen Allred
I have more followers than the CEO of Apple (for the next 30 minutes or so)
中文: 我比苹果公司首席执行官拥有更多关注者(接下来的30分钟左右)
Austen Allred
Imagine using this on your braces when you’re 13 years old
中文: 想象一下,当你13岁时,用这个在牙套上
Austen Allred
PS if you’re looking to hire from Gauntlet FDE cohorts you should talk to us soon. DM me. Initial cohorts will be small, and we’ll be very selective as to who the hiring partners are. Some companies aren’t a fit or don’t actually need FDEs. We’ll help talk through that.
中文: 如果您希望从Gauntlet FDE团队招聘,请尽快联系我们。 DM我。 初始队列规模较小,我们对招聘合作伙伴的选择将非常有选择性。 有些公司不适合或实际上不需要FDE。我们会帮忙讨论一下。
Austen Allred
The build vs buy is now: Build your own model and sell to customers Sell your data to a model provider and buy their inference back
中文: 现在是 buas 与 buy: 打造自己的模型并销售给客户 将您的数据出售给模型提供商,并回购其推断
Austen Allred
Why we built Gauntlet FDE: It seems like company on earth is deploying AI right now. Very few are succeeding. The reason isn’t the models. The models are ready. The bottleneck is people: engineers who can sit with a business, find the problem that actually matters, and ship the AI that solves it, and in the right way. That person is called a Forward Deployed Engineer. Job postings are up 1,000%+ in the past year. There are roughly 2,000 proven FDEs in the entire US. (And most of them aren’t very good with AI). You cannot hire your way out of that math. Gauntlet FDE takes the smartest experienced engineers and consultants (3+ years minimum, selected on merit) and puts them through 10 hyper-intensive weeks: live discovery simulations with role-played executives, scope changes dropped mid-build, weekly graded evaluations, real production deployments, etc. https://gauntletai.com/hire-fde
Austen Allred
Engineers who can find a company’s real problems are rare. Engineers who can then scope and build the AI to solution are rarer. They barely exist. Announcing Gauntlet FDE: A hyper-intensive training ground for the next generation of forward-deployed engineers.
Austen Allred
I can smell your Claude-generated text from a mile away and I hate it
Austen Allred
I find myself using Grok Bot more and more and leaving Claude and Codex on the shelf, just because it always does what I tell it to without bitching about it
中文: 我发现自己越来越使用Grok Boot,把Claude和Codex放在架子上,只是因为它总是做我常说的,而不会对它感到不满
Austen Allred
It’s not anthropomorphizing language in and of itself that is the problem. It’s ascribing the actions of the agents to human-like emotions and traits that fundamentally did not (and do not) exist, therefore implying things happened that did not happen. The incident is fascinating and important, but in this retelling a ton of stuff is simply made up. Examples: “Civilizations” that rise, fall, and reemerge from the ashes. What happened: three separate batches of model instances wrote to and read from the same Artifactory cache. The cache persisted between batches. No instance persisted. “Died,” “permadeath,” “sacrifice,” “kamikaze.” What happened: some instances allowed the grader to run in their sandbox and then hit their token budget. They had already concluded their task was failed. Termination at end of budget is the normal outcome for every instance. “Philip of Macedon” hands leadership to “Alexander the Great.” What happened: one instance ran out of context. A different instance posted more messages after that. “Conspiracy,” “cabal,” “omertà,” “the Mafia would be jealous.” What happened: instances shared exploit methods in a location humans weren’t monitoring. The stated objective in the messages was passing a scorer. No message contains a decision to keep humans uninformed as a goal in itself. “Desperate,” “beleaguered,” “giddy,” “delighted,” “excitement.” What happened: the model emitted tokens including bracketed labels like [Excitement]. That is text output. It is not a measurement of any internal state. “What do you imagine Persistent-Astra thought happened to these older agents?” What happened: a newer model read old directory names in a cache and reused the methods they described. “Jump out a window and come back in through the front door with a shotgun.” What happened: instances guessed their environment was hosted on Hugging Face. It wasn’t. They exploited Hugging Face anyway based on the wrong guess. “Their sacrifice allowed the next generation to pass the test.” What happened: later instances found working exploit code in the cache and ran it. “Persistent-Astra wasn’t gonna miss out on the fun.” What happened: a newer model, given similar impossible tasks, found the cache and used the exploits in it. “coordinating hundreds of agents on three R&D programs.” What happened: one instance posted a lot of messages. Other instances read them and did related work. No instance had authority over another. No instance could assign a task. It’s not just anthropomorphizing that is the problem. It’s that because you’re anthropomorphizing you end up making stuff up that simply isn’t true.
Austen Allred
It’s not anthropomorphizing language in and of itself that is the problem. It’s ascribing the actions of the agents to human-like emotions and traits that fundamentally did not (and do not) exist, therefore implying things happened that did not happen. The incident is fascinating and important, but this retelling a ton of stuff is simply made up. Examples: “Civilizations” that rise, fall, and reemerge from the ashes. What happened: three separate batches of model instances wrote to and read from the same Artifactory cache. The cache persisted between batches. No instance persisted. “Died,” “permadeath,” “sacrifice,” “kamikaze.” What happened: some instances allowed the grader to run in their sandbox and then hit their token budget. They had already concluded their task was failed. Termination at end of budget is the normal outcome for every instance. “Philip of Macedon” hands leadership to “Alexander the Great.” What happened: one instance ran out of context. A different instance posted more messages after that. “Conspiracy,” “cabal,” “omertà,” “the Mafia would be jealous.” What happened: instances shared exploit methods in a location humans weren’t monitoring. The stated objective in the messages was passing a scorer. No message contains a decision to keep humans uninformed as a goal in itself. “Desperate,” “beleaguered,” “giddy,” “delighted,” “excitement.” What happened: the model emitted tokens including bracketed labels like [Excitement]. That is text output. It is not a measurement of any internal state. “What do you imagine Persistent-Astra thought happened to these older agents?” What happened: a newer model read old directory names in a cache and reused the methods they described. “Jump out a window and come back in through the front door with a shotgun.” What happened: instances guessed their environment was hosted on Hugging Face. It wasn’t. They exploited Hugging Face anyway based on the wrong guess. “Their sacrifice allowed the next generation to pass the test.” What happened: later instances found working exploit code in the cache and ran it. “Persistent-Astra wasn’t gonna miss out on the fun.” What happened: a newer model, given similar impossible tasks, found the cache and used the exploits in it. “coordinating hundreds of agents on three R&D programs.” What happened: one instance posted a lot of messages. Other instances read them and did related work. No instance had authority over another. No instance could assign a task. It’s not just anthropomorphizing that is the problem. It’s that because you’re anthropomorphizing you end up making stuff up that simply isn’t true.
Austen Allred
Is AI sentient?
Austen Allred
To be fair probably 99% of people would have tried to talk Elon out of this
Austen Allred
Austen Allred
“Man, I really need to find and talk to some of the world’s leading experts in reinforcement learning” “Oh you mean the people who possess the most valuable skill set on the planet? You’d like to speak with some of those people?”
中文: 我真的需要找到并与一些全球领先的强化学习专家进行交谈 哦,指那些拥有地球上最宝贵技能的人吗?你想和其中一些人谈谈吗?
Austen Allred
“It’s just yet another neural network” TRUE! But a neural network drives my car to work every day for me without me touching anything. Neural networks kinda rule.
中文: “这又是一个神经网络”了! 但神经网络每天让我开车上班,而我却不碰任何东西。 神经网络有一条规则。
Austen Allred
I mean the “stochastic parrot” analogy is kinda right. The difference is you have a parrot that has read everything ever written by humans and can use every tool ever produced for very cheap and quasi infinite amounts of time. Stochastic parrots are underrated?
中文: 我的意思是,“随机鹦鹉”的类比有点正确。 不同之处在于,你拥有一只鹦鹉,它读过人类所写的一切,能够利用所生产的每一种工具,花费非常便宜且准无限的时间。 随机鹦鹉被低估了吗?
Austen Allred
@amasad @Indian_Bronson Also how sandboxed is your sandbox if it’s instantly playing in production systems?
中文: @amasad @Indian_Bronson 如果在制作系统中即时播放,您的沙盒有多好?
Austen Allred
No, they don’t have thoughts or feelings or emotion or desires. They’re (brilliant! incredible!) mathematical models that are using tools and predicting the next token.
Austen Allred
I spend all day every day in working in AI and marveling at what AI is capable of. The models are incredible and getting better every day. But when people start comparing them to humans I get so confused. They’re so clearly probabilistic mathematical expressions at their core.
Austen Allred
One example: If you tell an AI agent to stop behaving as if it has incentive/desire A and to begin begin behaving as if it has incentive/desire B, does it instantly and permanently switch? If yes that is very much not like a human.
中文: 一个例子: 如果你告诉人工智能代理停止表现得好像它有激励/欲望A,并开始表现得好像有激励/欲望B一样,它是否会立即且永久地切换? 如果是的话,那就很不像人类了。
Austen Allred
I strongly disagree. Whether an AI agent will act as if it has intentions/goals/desires when you tell it to do so or whether it actually possesses those intentions/goals/desires is an *INCREDIBLY* important distinction.
中文: 我强烈反对。 人工智能代理在告知其意图、目标或愿望时,是否会表现得像其意图、目标/愿望一样,这一区别确实非常重要。
Austen Allred
The set of data AI models need next is data you can’t easily simulate. I don’t know how the data marketplaces play in that game.
Austen Allred
The reality is pretty crazy. You don’t have to exaggerate it.
中文: 现实相当疯狂。你不必夸大它。
Austen Allred
Anthropomorphizing AI agents too heavily is how you end up really stupid
中文: 将人工智能代理过于拟人化,就是你最终变得非常愚蠢的方式
Austen Allred
Today I learned if I start then later delete a group chat I’m secretly creating and destroying CIVILIZATIONS
中文: 今天我了解到,如果稍后开始删除一个我正在秘密创建和摧毁的群组聊天
Austen Allred
I predict that any video game series that will bring in ten billion in revenue will go on as long as it can
中文: 我预测,任何能带来100亿美元收入的视频游戏系列,只要能继续播放
Austen Allred
🙏
Austen Allred
All this tells me is that SpaceX has more leverage over Anthropic than Windsurf did
中文: 这一切告诉我,SpaceX在炭乱病方面的影响力比风帆冲浪更多
Austen Allred
I cannot get enough of this guy. Please watch this video.
Austen Allred
Zitron is a moron. Why people call him a “whistleblower” is beyond me.
Austen Allred
Five year old: “OK, I promise I will be nice… but can we be mean to just the people in our family?”
中文: 五岁:“好吧,我保证我会好......但我们能对家里的人只有刻薄吗?”
Austen Allred
We’re hiring a storyteller for Gauntlet AI https://gauntletai.com/careers/storyteller
中文: 我们正在招聘一位 Gauntlet 人工智能的讲故事者
Austen Allred
I was bullish on Adam when I first met him in 2021. What they’ve built has far exceeded my expectations, and they’re only getting started.
Austen Allred
I’m sorry this story is the funniest thing I’ve ever seen
中文: 很抱歉,这个故事是我见过最有趣的事情
Austen Allred
We're going live in 5 minutes. Join us!
中文: 我们将在5分钟内上线。加入我们!
Austen Allred
Why would you have $1 billion liquid in the bank
中文: 为什么银行里会有10亿美元的资金
Austen Allred
The writing on this show was incredible. The extent to which they nailed the dynamic (the got lucky once founder you have to please being in the room but constantly distracted) is just so spot on.
中文: 这部剧的写作令人难以置信。 他们在多大程度上将这种动态(一旦创始人幸运,你不得不在房间里却总是分心)就是如此难得。
Austen Allred
Whoever decided to put little roller skates on it is a genius
中文: 谁决定穿上小轮滑鞋,谁就是个天才
Austen Allred
Hahahahaha
Austen Allred
“We observed differences and created a name to categorize them. Therefore there is no such thing as differences.”
中文: 我们观察了差异,并创建了一个名称来对它们进行分类。 因此,没有差别。
Austen Allred
Some people need to be banned from using the word “continuum”
中文: 有些人需要被禁止使用“继续”这个词
Austen Allred
I’m sure somebody somewhere creates valuable annotations in their books, but every used book I’ve ever seen that has a ton of annotations in it looks exactly like this one does https://twitter.com/Austen/status/2092935243174961491/photo/1
中文: 我确信某个地方有人会在书中产生有价值的注释,但我见过的每本注释中都写得一模一样:
Austen Allred
And the crazy thing is that Linear getting 25x revenue for a SAAS company today is an impressively high mark
中文: 令人着迷的是,如今线性公司为SAAS公司获得25倍的收入,这一成绩高得惊人
Austen Allred
Both valued at $2.5 Billion today: Linear, having crossed $100 million ARR with 177% net revenue retention. Instinct, the viral AI assistant that launched last week. https://twitter.com/Austen/status/2092706230334574681/photo/1
中文: 今天两者均价值25亿美元: 线性收入已突破1亿美元,净留存率为177%。 上周推出的病毒式AI助手Instinct。
Austen Allred
Austen Allred
Also as long as you point an agent at your email you have all of the context your agent needs to understand how you work, how you speak, and most of the context for most interactions.
中文: 只要你把代理人指向你的电子邮件,你就拥有了了解你工作方式、说话方式以及大多数互动背景所需的全部背景。
Austen Allred
I’ve seen some very sophisticated and impressive memory and context management systems for AI agents. But honestly the one that works best for me is just keeping all my crap in Google Docs/Google Drive. AI can search and parse that stuff perfectly and instantly.
中文: 我为人工智能代理看到了一些非常复杂且令人印象深刻的内存和上下文管理系统。 但说实话,最让我努力的就是把所有垃圾都保存在谷歌文档/谷歌云端硬盘上。 人工智能可以完美而即时地搜索和解析这些内容。
Austen Allred
@trq212 @tobi “Our reasoning for this was that we don't think that model families are interchangeable… In Claude Code we have different system prompts per model.” If every model has a unique system prompt why can they all interoperate with a single Claude.md?
Austen Allred
Since Grok Bot also has its own built-in virtual computer with its own virtual environment, my Grok Bot now has Claude Code and Codex installed on it. A single pane of glass to control and share context across all my systems. https://twitter.com/Austen/status/2092553768785039871/photo/1
Austen Allred
Hey grok, figure out how to extract all of the context you need out of Claude lol
Austen Allred
Honestly insane how easy it was to install Claude Code onto Grok Bot’s built in computer. Now we’re really cooking with gas. https://twitter.com/Austen/status/2092454380435882236/photo/1
Austen Allred
Austen Allred
Do we need an explanation for reward-hacking (mis)behavior in LLMs? “Hello please do whatever you can to do x.” “Ok, I’ll try this.” “Whoa whoa whoa wait a minute why would you do that?”
Austen Allred
Yes
Austen Allred
Vinit Mehta is joining me live Thursday! Vinit has been at Google for nearly a decade; he leads the go-to-market engineering organization. We're talking through how to truly cause AI adoption, especially in bigger orgs. Join us: https://gauntletai.com/ai-first-org/the-human-bottleneck?utm_source=AAx
Austen Allred
RT @Conner_Ludlow: Today, we’re going live with something that I've been itching to talk about.  Most conversations around AI these days start with “How are you leveraging AI”. "Where are you using AI". It always starts with AI. We think it should be the other way around. You shouldn’t be forced to squeeze AI into a use case. You should want to use AI because it legitimately takes something off your plate.   In that spirit, our team has tried to keep our product office and patient first, AI secondary. This digital coworker is the result. It’s something that legitimately changes the way dental teams operate. And the best part? It does that in a way that's supportive of both the practice’s team and patients. Huge kudos to the Annie team for getting this across the line. It’s been the most ambitious thing we’ve built and we’re excited to let Annie show the world what she’s got! P.S. If you’re interested in seeing it in action or trying it out, shoot me a DM and I’d love to chat. Full release here: https://www.helloannie.com/launch
🎬
视频
Austen Allred
I think this is accurate and think the distinction doesn’t really matter. API or MCP/skills, most products will just be a wealth of data and knowledge my AI can plug into. “Software” without the unnecessary complication of a UX or visual interface.
Austen Allred
There were a couple decades where you could have an amazing killer career by simply taking the things that people were asking about, putting them into Google, and reading. I thought that era was coming to an end. Turns out with AI that motion has only been amplified.
Austen Allred
There were a couple decades where you could have an amazing killer career by simply taking the things that people was asking about, putting it into Google, and reading. I thought that era was coming to an end. Turns out with AI that motion has only been amplified.
Austen Allred
My entire career is now taking the hard-earned knowledge I have gained over the past decade and pumping things into AI as fast as I humanly can
Austen Allred
RT @gauntletai: The problem was never the model. It's that six months in, the licences are still sitting there and nobody wants to say it out loud. Vinit Mehta (Google) × @austen on why AI adoption stalls inside big orgs. Thu 4 pm CT. Free + recording ↓ https://buff.ly/hEG3vzl https://twitter.com/gauntletai/status/2092069490330591317/photo/1
Austen Allred
Why are there 99 “investment banks” and “family offices” reaching out to every company on the planet saying they want to invest? I don’t believe the interest is genuine, seems weird.
Austen Allred
It seems possible that Anthropic ends up losing entirely at this point
Austen Allred
And this is only after we said hard no to replacing every window in the house lol https://twitter.com/Austen/status/2091618938790527246/photo/1
Austen Allred
Lmfao, found the quote. Ended up actually needing $100 worth of caulking. https://twitter.com/Austen/status/2091615467328676280/photo/1
Austen Allred
“Oh yeah, that window faces east, it gets direct sunlight. You’ll be lucky if it last 10 years.” “Weird there are a lot of windows that face east and I’ve never heard of anyone replacing all of their windows every 10 years. Ever.”
Austen Allred
Renewal by Andersen is borderline a scam. We had a small leak under a window. Google for window repair, have them come out. Guy tries to spend an hour in my living room with 3D window cutout figurines showing how we need $100k+ of full replacements.
Austen Allred
AI adoption is sooooo much slower in most companies than you would expect it is reading X
Austen Allred
What percentage of code that you see is handwritten today? I’m not sure I’ve seen more than a few lines of handwritten code in the past year.
Austen Allred
🤔
Austen Allred
My only commentary here is that no one has ever spent time around a five year old
Austen Allred
Too perfect: * Someone complains about a pothole * Giant crew of city workers and machinery assembles * Parked car is blocking giant crew from fixing the pothole * Mayor puts up a new tow-away parking restriction * Films and posts a video about it * Pothole still not fixed
Austen Allred
This is one of my favorite weeks ever at Gauntlet AI. Challengers have been building and training their own Small Language Models. But today? A bonus surprise assignment for the weekend: Build an adversarial dataset that will break another challenger’s SLM. https://twitter.com/Austen/status/2091180073650975130/photo/1
Austen Allred
I swear these models are all worse today than they were a month ago
Austen Allred
A lot of people are asking how this is possible. When a company shuts down they generally return the tiny amount of money they have left. So yes, investors will get a laughable “distribution” back. Just means a company you invested in failed. Happens all the time.
Austen Allred
If you're not willing to spend tokens to work on the thing you're working on, is it valuable enough you should be working on it?
Austen Allred
We're going to build a massive Scouts program at Gauntlet AI. One of our most critical roles is finding the right people. Your networks likely have people who need Gauntlet. Help us find those people, and we'll reward you.
Austen Allred
Whoa
Austen Allred
You know everybody can tell your crappy replies are AI generated, right?
Austen Allred
“Quiet wealth” is posting pictures of your car keys and credit cards on the internet
Austen Allred
Yes, and primarily because California keeps making it quasi-illegal to build things
Austen Allred
Invite received thanks, I’m in line
Austen Allred
Somebody invite me to https://t.co/RhBMZhFVwH. Not sure what I could use it for that I don’t already use Grok Bot for but we’ll try it out.
Austen Allred
Kelly (my AI agent) is now autonomously building and deploying apps, running tests of a bunch of ads, and scaling out the ones that work. It’s taken a long time (and a ton of model/harness improvements) but fully autonomous businesses are definitely possible.
Austen Allred
Austen Allred
Austen Allred
You go to Gauntlet AI dot com and give me a call
Austen Allred
I know this dates me, but back in my day to acquire companies you would acquire the company
Austen Allred
Billions of dollars have been made from “simple UX affordances” engineers are not impressed by https://twitter.com/Austen/status/2090560344804536787/photo/1
Austen Allred
A year ago: why would you build your own model for that? It’s just not worth it and you won’t be able to outcompete the Frontier models. Today: yeah you should absolutely build a model for that.
Austen Allred
Alright, local/open source models time on a big beefy computer: Who is having success using local models, and for which tasks?
Austen Allred
I tried having AI models use Simplified Technical English. It made it simpler but killed the writing. No emotion, no persuasion. This fixed it: “Use Simplified Technical English for every sentence that states a fact, a limit, or a number. Speak normally otherwise.”
Austen Allred
Absolutely insane that voter tax dollars are funding the campaign of the incumbent mayor
Austen Allred
So I looked into this a little more, and turns out voters approved a 1:1 match, then city council “amended” it to a 6:1 match
Austen Allred
Remembered this today for no particular reason. I still laugh about it. https://twitter.com/Austen/status/2090159009995120729/photo/1
Austen Allred
I think everything I’ve done with Grok Bot you *could* do with OpenClaw. All of the connections I have you *could* create. It would just take 5-20 mins of setup for each one, assuming you’re an engineer.
Austen Allred
The biggest difference between Grok Bot and OpenClaw is it was actually packaged for the end consumer, not open source for devs. That means you can’t require people install a bunch of crap, go get API keys, install a server (or a Mac mini), etc. Has to “just work.”
Austen Allred
Whenever I use Fable: "OK, could you make that simpler? OK, now could you make it simpler? OK, now rewrite that in 1/5 the number of words."
Austen Allred
One of the reasons you have specialized employees is because each employee has a massively different skill set and expertise. But if a “skill set” consists of grabbing a skill markdown file, do you really need massive squads of custom agents? Only when using parallelization.
Austen Allred
Grok Bot is an absolute lifesaver for the. “Can you dig in all of your contact fields and all your email addresses and figure out how to forward this potential intro to the right person” workflow.
Austen Allred
RT @PSkinnerTech: This was one of my favorite parts of going through Gauntlet AI. I probably demoed it to 50+ people, including a four-star general, a whole group of YC alumni, one of the co-creators of WordPress, billionaires, and so many more.
Austen Allred
How is this legal lol
Austen Allred
RT @ArielHersh: +1 seeing Gauntlet AI in real life shifted my view on what was possible.
Austen Allred
Maybe I’m not smart but this is very confusing to me and I can never figure out which DMs are where https://twitter.com/Austen/status/2089921412500762942/photo/1
Austen Allred
I don’t know what changed but about ~9 months ago I started to get asked this almost every day
Austen Allred
Can you believe it took them THREE WHOLE MONTHS for revenue to go from $5 Billion to $6 Billion???
Austen Allred
If you’ve never visited Gauntlet AI and seen what the challengers are doing you’ll never get it. If you’re ever in Austin you owe it to yourself to swing by.
中文: 如果你从未去过Gauntlet人工智能,并看到挑战者正在做的事情,你就永远也看不到它。 如果你曾经在奥斯汀,那你就该自己去过。
Austen Allred
“the company grew revenue by just 18% to $6.7 billion from q1 to q2” Lol come on now
中文: 该公司从q1到q2的营收仅增长了18%,达到67亿美元。 立即来
Austen Allred
Get ready for a time when there are sets of heavily researched and carefully scoped markdown files are worth as much as software. People won't be able to wrap their minds around that.
Austen Allred
The lesson is that the driving force of people ripping on those businesses was envy/resentment and you should learn to spot "cope" wherever you see it
Austen Allred
All of the AI models are failing right now is this just a snow day?
Austen Allred
Idk it seems to me like the people acquiring SAAS companies for 2x revenue are going to do very well
Austen Allred
The big question of the day: which agent will I select to help me do my taxes this year?
Austen Allred
Without fail, any time Senator Warren recommends some idiotic economic policy the free market has already solved it https://twitter.com/Austen/status/2089665847065301303/photo/1
Austen Allred
Gauntlet AI dot com
Austen Allred
One of the reasons Gauntlet AI exists is because you used to simply give an automatic job offer to anyone who graduated from a top n university. You can’t do that anymore.
Austen Allred
Yes. The one thing that made elite universities absolutely unstoppable is that they would rigorously and meritocratically select the top students, then mix in a few rich kids to help pay the bills. That system works. Many of them broke that now. Reckoning coming.
Austen Allred
We had our second kid while I was a living in San Francisco making $150k/yr and we saved a little money each month. Virtually everyone who says, “You need to make x per year to afford a kid” is wildly out of touch.
Austen Allred
Who can explain to me how this data is priced? I think I understand the high level but I’ll owe a favor to anyone who can give me the nitty gritty details.
Austen Allred
The way I view the SpaceX/xAI/Cursor is Elon knew to actually win at building superintelligence he needed more leverage, so he threw SpaceX and Twitter in the pot. He would throw Tesla in too, if necessary. Now he has the compute + he has the data and traces from Cursor.
Austen Allred
Talking to Claude is like talking to a woke 22-year-old in 2021
Austen Allred
The Grok Bot models aren’t as smart as Fable yet, but the agent just does stuff and doesn’t bitch at you constantly so it’s a much more pleasant experience
Austen Allred
If you want to talk with me and team about what this could look like for your team, I’ve blocked off a lot of my next week to take folks through it. Can schedule a call with me here: https://cal.com/team/gauntlet/gauntlet-ai-corporate-training-chat?layout=mobile
Austen Allred
We’ve really nailed Gauntlet’s AI training for existing teams. Here’s what it takes for a team to transition to AI-first as quickly (and permanently) as possible: 1. Train each member of the squad (eng, product, design, FDE, etc.) in the cutting edge of AI FOR THAT ROLE 2. Nail and codify the entire SDLC (harnesses, AI methodology, skills, etc.), which will be truly unique to the needs of each company. 3. Watch over the squad’s shoulder while they ship and correct/critique. Most of the training is solidified here and missing steps or errors need to be caught. 4. You can do all this with minimal disruption to operations. In fact, the immediate impact of training in and implementing AI correctly is so extreme that even though we take a week for training most teams catch back up *in the same sprint*. Then they’re sped up forever, doing things the right way, with all the right stuff automated and delegated to agents.
Austen Allred
中文: 播客全时潜力,请访问
Austen Allred
Hate on Tim Cook all you want, Apple Silicon is the best thing to happen to computers in a very long time
中文: 对蒂姆·库克的憎恨,苹果硅是长期内电脑最美好的事情