背景: OpenRouter 是一个平台,为开发者提供单一 API 来访问和切换众多 AI 模型(例如来自 OpenAI、Anthropic 等公司的模型),抽象了处理多个提供商、定价和集成的复杂性。Stripe 是一家领先的全球金融基础设施公司,以其在线支付处理 API 闻名,但一直在向邻近领域扩展,如资金管理、公司卡,以及现在的 AI 基础设施。
智谱 AI 正式上线了 GLM-5.3 大语言模型的 API 服务,该模型在 AA 综合智能指数中取得 60 分,与 Kimi K3 并列开源模型第一,并与 Claude Fable 5、GPT-5.6 Sol 等闭源旗舰模型同级。模型权重将于下周五开源,其 API 定价与 GLM-5.2 持平。 此次发布意义重大,因为它声称能以更低的成本提供与领先闭源模型相当的性能,有望为开发者和企业降低使用前沿智能的门槛。作为一个在关键基准测试中达到顶级水平的开源模型,它可能会加剧大语言模型市场的竞争,并推动高性价比 AI 的普及。 GLM-5.3 并未更换基座模型,而是在 GLM-5.2 的基础上,通过在长周期环境中进行一个月的强化学习,使其编码能力提升了 50%。它是一个总参数量约 743B 的混合专家模型,但每次推理仅激活约 40B 参数,这有助于降低其调用成本。
一篇技术文章详细介绍了如何通过分析照片中山脉的阴影来确定太阳位置,进而利用 CUDA 加速的暴力搜索算法,将地形与数字高程模型(DEM)进行匹配,从而解决了一项地理定位挑战。作者成功确定了岛屿的坐标、度假村的名称以及相机的朝向。 这展示了一种新颖且实用的开源情报(OSINT)应用,它结合了计算机视觉、天文学和高性能计算,超越了简单的图像匹配。它展示了 GPU 加速如何使计算密集型的地理空间分析对个人研究者变得可行,并对自动驾驶导航和法医调查等领域具有启示意义。 该方法涉及根据阴影几何计算太阳的方位角和高度角,然后使用 CUDA 内核高效地测试全球数字高程模型中数百万个潜在位置,以寻找匹配的地形轮廓。一个关键限制是要求照片清晰、阴影分明且地形特征可识别,同时精度取决于可用高程数据的分辨率。
在最近的一期播客中,开发者 Simon Willison 提出了一个细致的论点:在使用 AI 编程助手时,测量代码行数可以成为衡量生产力的一个有意义的指标,这直接挑战了软件工程中长期存在的一种信念。他认为,虽然一名人类工程师每天可能产出 50-200 行经过调试、可用于生产的代码,但熟练使用 AI 助手可以显著提高这一产出,只要代码质量得以保持,就代表着真实的生产力提升。 这一点很重要,因为它重新构建了 AI 时代关于软件指标的核心辩论,表明当生产机制从纯粹的人力转向人机协作时,传统上对代码行数作为指标的否定可能需要重新评估。它强调,对于工程团队来说,新的瓶颈可能不再是代码输出速度,而是在代码库现在可以呈数量级更快增长的情况下,维持其概念完整性所需的认知能力。 Willison 强调,只有当 AI 生成的代码保持质量——即可维护、经过测试且可用于生产时,这种生产力提升才成立。他还提出了一个关键挑战:当 AI 助手使得添加功能变得廉价且容易时,如何保持’概念完整性’(一种连贯、统一的设计),否则可能导致代码库变成一个支离破碎的’温彻斯特神秘屋’。
rss · Simon Willison · 8月19日 22:46
核验: 多源印证
背景: ‘概念完整性’这一术语源于 Frederick P. Brooks 的开创性著作《人月神话》。它指的是软件系统的一个特性,即其核心概念作为一个流畅、连贯的整体协同工作,这被认为是良好设计和可维护性的关键。AI 编程助手是先进的 AI 工具,可以自主规划、执行和验证多文件的代码变更,超越了简单的代码补全,更像自动化的编程伙伴。
背景: Anthropic 的 Astra 是一款未发布的前沿 AI 模型,据报道已解决多个高难度数学问题,表明其具备高级推理能力。“关键网络安全能力阈值”这一概念指的是 AI 能力达到一定水平,模型可以自主执行或极大助力复杂的网络操作,从而构成严重的双重用途风险。工作负载隔离和网络隔离等安全措施是用于分割和控制 AI 系统的技术,旨在防止未经授权的访问或数据泄露。
背景: MLX 是苹果推出的一个数组框架,专为在 Apple Silicon 上进行高效的机器学习而设计,它利用 Metal API 进行 GPU 加速。视频扩散变换器 (DiT) 是一种生成模型架构,它在扩散过程中使用变换器来创建时间上连贯的视频。DMD(分布匹配蒸馏)是一种蒸馏技术,能大幅减少扩散模型所需的采样步数,从而实现更快(通常是单步)的生成。
背景: Go 语言在 1.18 版本中引入了泛型,允许使用类型参数编写函数和类型。然而,泛型类型上的方法本身不能有额外的类型参数,这一限制在 Go 1.27 中得到了解决。后量子密码学指的是设计用于抵御经典计算机和量子计算机攻击的算法,NIST 正在领导其标准化工作。UUID(通用唯一标识符)是用于唯一标识信息的 128 位数字,google/uuid 包一直是 Go 社区的事实标准。
Simon Willison 让在 Claude Code for web 中运行的 Claude Fable 5 执行一项研究任务:评估 smolvm 沙箱,以安全地运行不受信任的 Python 和 JavaScript 代码。当 Claude Code 环境缺少硬件虚拟化支持时,该 AI 自主设计并执行了一个使用 GitHub Actions 来运行测试的解决方案。 这展示了一种解决 AI 智能体开发中关键挑战的实用方法:安全地执行可能有害或资源密集的用户提供代码。AI 主动解决问题的过程,也突显了大语言模型在处理涉及环境限制的复杂、多步骤技术工作流方面日益增强的能力。 Claude Code for web 环境缺少 /dev/kvm 设备和 CPU 虚拟化标志,无法直接使用需要 KVM 的 smolvm。AI 的解决方案是编写一个 GitHub Actions 工作流,在拥有 KVM 访问权限的运行器上安装 smolvm 并执行测试脚本,从而成功测试了沙箱的功能。
rss · Simon Willison · 8月19日 23:16
核验: 多源印证
背景: Smolvm 是一个开源的沙箱基础设施,它使用微虚拟机(如 Firecracker)来提供快速、隔离的代码运行环境,启动时间不到 200 毫秒。它专为安全执行 AI 生成或用户提交的代码等场景设计,可施加资源限制和网络隔离。Claude Code 是一个由 Anthropic 的 Claude 模型驱动的基于终端的编码环境,有时会限制系统访问,如此处所见,它缺乏嵌套虚拟化支持。
Jeremy Morrell 发表了一个假设,即大语言模型(LLMs)与现代沙盒技术共同为构建 Web 上的可扩展软件创造了新的机遇。他认为,LLMs 极大地降低了编写扩展的成本,而沙盒技术则降低了部署成本并提供了强大的安全边界,使得应用程序可以拥有一个安全的核心,同时允许用户安全地进行扩展。 这一点很重要,因为它可能实现软件定制的大众化,将权力从开发者转移到最终用户,从而在不牺牲安全性的前提下实现高度个性化的应用程序。这代表了软件架构的一个重大转变,可能催生出跨行业、更具适应性且以用户为中心的 Web 平台。 该假设特别关注 Web 作为平台,并利用基于浏览器的沙盒机制来保障安全。一个关键的技术前提是将用于代码生成的 LLMs 与用于安全执行的沙盒技术相结合,共同应对传统可扩展系统中开发成本高和安全风险大的挑战。
rss · Simon Willison · 8月19日 22:56
核验: 多源印证
背景: 可扩展软件的设计允许在不改变其核心架构的情况下添加新功能或进行修改,通常通过插件或 API 实现。现代沙盒是一种安全机制,用于隔离运行中的程序,在 Web 浏览器中常用于限制网页代码,防止其损害用户系统。像 GPT-4 这样的大语言模型(LLMs)是能够生成和理解文本(包括代码)的 AI 系统,可以自动化那些以前需要大量人类专业知识才能完成的任务。
Sports are amazing environments to learn. When you play a sport for thousands of hours, you start to see the world through that sport. It is a simple fact—your biological neural network is being conditioned to respond to the behavior incentivized by the rules of the sport.
The funny thing is that most people choose their sports for accidental reasons such as parents, geography, or school programs. People rarely think about how the particular sport you play influences how your brain thinks more generally. Going a step further, playing the right sport may even benefit your career.
My two favorite sports are tennis and soccer. Tennis is one of the best sports for teaching consistency. In tennis, there are hundreds of points in a match, and each point is worth exactly one unit, regardless of whether your opponent made an unforced error or if you constructed the most beautiful point ending with a winner. Tennis is low-variance optimization—you win by reducing unforced errors, playing percentages, and grinding out small advantages. Tennis is also an individual sport, which teaches you to rely on yourself consistently.
Tennis has a similar cognitive reward shape to professions like being a surgeon or a pilot. Surgery and aviation require consistency, self-accountability, and deep focus. And similar to how you can only win one point at a time in tennis no matter how spectacular it was, there is no extra credit for the best appendectomy or the smoothest SFO-JFK flight. Your craft is to provide consistency with very low tolerance for error.
On the other hand, the tennis mindset transfers relatively little to entrepreneurship. Entrepreneurship is a high-variance, team game where failure is tolerated and occasional creativity gets rewarded exponentially. Minimizing unforced errors in tennis is a totally different mindset from deciding whether to make a moonshot business move that will likely fail but could potentially net a billion dollars. Obviously I am not saying that tennis players cannot be great entrepreneurs, but I do think it is a totally different cognitive reward shape.
Being a forward in soccer has a much closer reward shape for entrepreneurship. What a forward in soccer learns is to create many small chances. It is a fact that most of the game, you are not scoring—even if you look at all the times that Mbappe got on the ball in one of his best games, most of those led to nothing! But all that matters is creating enough chances to score once (or a few times) and win the game. If you break down a 90-minute game for a forward, almost all the time is failure or noise, a few minutes will be leverage, and a few seconds will determine the fate of the game. I have not played soccer for thousands of hours, but I can imagine that being a lifetime forward in soccer would teach you to be comfortable with failure and asymmetric returns.
In summary, I am claiming that there can be substantial value when the cognitive reward shape of your sport mirrors that of your career. I’ll admit that I’ve done some cherry-picking for illustration purposes—entrepreneurship also requires consistency and error avoidance; and goalies in soccer have reward shapes that are very different from strikers. But I think the point stands. If sports shape how we perceive risk, effort, and reward, then we should choose them wisely.
Follow Builders · X 动态 · Peter Yang · 8月19日 04:25 UTC · 喜欢 20 · 转发 0 · 回复 13
The author proposes building an app or AI agent to track a streak of days without bringing a phone to the bedroom, noting its positive impact on sleep discipline.
中文摘要作者提议开发一个应用或 AI 智能体,用于追踪不带手机进卧室的天数,并指出这对其睡眠纪律有积极影响。
Aaron Levie argues that the value created by integrating AI models into end-user workflows is greater than expected, but successfully deploying AI agents in critical enterprise processes requires tailored product experiences and data integration.
中文摘要Aaron Levie 认为,将 AI 模型集成到最终用户工作流中所创造的价值超出预期,但要在关键企业流程中成功部署 AI 智能体,需要量身定制的产品体验和数据集成。
Follow Builders · X 动态 · Madhu Guru · 8月19日 03:31 UTC · 喜欢 39 · 转发 2 · 回复 9
The post advocates for prioritizing high-quality evaluation to establish a trusted performance baseline for AI products before optimizing for cost efficiency.
中文标题Peter Yang: 3. AI 叠加在现有工作之上,而非取代它 😭 团队现在花了更多时间...
Follow Builders · X 动态 · Peter Yang · 8月19日 00:48 UTC · 喜欢 3 · 转发 0 · 回复 0
The author observes that AI tools are adding to existing workloads rather than replacing them, as teams spend more time managing AI agents without reducing their other tasks.
中文摘要作者观察到,AI 工具正在增加现有工作量而非取代它,因为团队花了更多时间管理 AI 智能体,却没有减少其他任务。
Follow Builders · X 动态 · Peter Yang · 8月19日 00:48 UTC · 喜欢 2 · 转发 0 · 回复 1
Data shows a significant increase in code contributions from non-engineers like product managers and designers over two years, indicating a trend toward broader participation in software development.