Wednesday, July 15, 2026
Agents are rewriting the stack from the inside out.
July 15 · 10 videos
Bun rewrote 535,000 lines of code in 11 days.
They used 64 parallel Claude agents.
It cost $165,000 in tokens.
OpenAI ignored a government ban to ship GPT-5.6.
Cursor is automating its own researchers.
The human bottleneck is finally breaking.
“Rewrites are now good. The worst thing you could do is now actually fine.”
Computer-Use 2.0: Agents Just Got Multi-Cursor — Francesco Bonacci, Cua
Francesco Bonacci · AI Engineer · 16 min
Watch on YouTube →Francesco Bonacci explains the shift from screen-hijacking agents to background system-level drivers. This transition aims to increase reliability and reduce token costs in professional workflows.
- Agent pass rates improved from 62% to 80% on 4K benchmarks using the cua driver.
- Token consumption dropped by 34% by focusing on specific window states instead of full screens.
- Current models show a 0% success rate when designing electrical schematics from scratch in KiCad.
- The cua driver uses undocumented accessibility APIs like UI Automation on Windows and AX on macOS.
- Demand-based warm pools for 40GB containers prevent GPU idle time during heavy environment startups.
- Trust remains the primary barrier to adoption, requiring verifiable benchmarks like CUABench.
NEW: Alex Hormozi Answers Your Questions on Reddit
Alex Hormozi · Alex Hormozi · 89 min
Watch on YouTube →Alex Hormozi breaks down the mechanics of scaling a $250M portfolio by identifying system bottlenecks. He emphasizes leveraging existing high-level skills over starting new ventures from zero.
- Business growth is a function of identifying and removing the single active rate limiter.
- The 5-11-20 rule suggests the top 20% of an audience has 10-11x the wealth, justifying 5x price jumps.
- Intelligence is defined by the speed of behavioral change following new information or trauma.
- Action is the primary antidote to anxiety, while vague fears are solved by defining worst-case outcomes.
- B2B services should stay close to the money by directly increasing revenue or decreasing costs.
- The all-cash exit for Gym Launch and Prestige Labs was valued at 46.2M.
Recursive Model Improvement — Lee Robinson, Cursor, SpaceXAI
Lee Robinson · AI Engineer · 20 min
Watch on YouTube →Lee Robinson details how Cursor uses smarter models to automate the training and evaluation of future versions. This recursive loop aims to remove human researchers as the primary bottleneck.
- Cursor is moving toward full pre-training on custom infrastructure to increase model efficiency.
- The Colossus supercomputer was built in 122 days and now houses 200,000 GPUs.
- Teacher-Student Textual Feedback is used as a novel RL method to solve credit assignment in long tasks.
- The half-life of an evaluation decreases as models improve, requiring constant retirement of benchmarks.
- Vertical integration from the IDE to custom silicon via Terafab provides a significant competitive edge.
- Revenue is shifting from simple chat interfaces to complex agentic usage patterns.
The most controversial rewrite in history just shipped...
Fireship · Fireship · 5 min
Watch on YouTube →Fireship covers Bun's decision to rewrite its entire codebase from Zig to Rust using AI agents. The move challenges the traditional industry dogma that full rewrites are strategic suicide.
- 64 parallel Claude agents completed the 535,000-line rewrite in just 11 days.
- The project cost approximately $165,000 in AI tokens and compute resources.
- The new Rust binary is 20% smaller and resolved 128 long-standing bugs.
- AI agents achieved a peak output of 1300 lines of code per minute during the port.
- Language choice is now influenced by AI-assisted developer velocity and training data availability.
- The transition sparked a public dispute between Bun founder Jared Sumner and Zig creator Andrew Kelly.
Context engineering with Dex Horthy
Dex Horthy · The Pragmatic Engineer · 93 min
Watch on YouTube →Dex Horthy argues that fully automated dark factories lead to codebase collapse within months. He advocates for context engineering to maximize model leverage within the reliable smart zone.
- Codebases become unfixable after 3 to 6 months if humans stop reading and reviewing the code.
- The smart zone for frontier models is typically the first 100k to 200k tokens of context.
- Frontier LLMs generally fail to follow more than 150 to 250 specific instructions reliably.
- Slow loops use cron-triggered agents to fix one specific code anti-pattern per night.
- The RPI framework (Research, Plan, Implement) keeps humans focused on high-leverage architectural design.
- Using a smarter model is usually more cost-effective than human hours spent on context engineering.
Simon Willison in conversation with Cat Wu & Thariq Shihipar, Anthropic
Cat Wu · AI Engineer · 51 min
Watch on YouTube →Anthropic engineers discuss a paradigm shift where manual implementation is no longer the primary engineering bottleneck. They highlight how internal agents now land the majority of their product PRs.
- The Claude Tag agent now lands 65% of internal product PRs at Anthropic.
- System prompts for frontier models have been reduced by 80% as internal model judgment improves.
- The timeline from initial idea to production build has compressed to as little as one week.
- Auto Mode uses a Sonnet classifier to judge the security of tool calls in real-time.
- Developer value is shifting from writing lines of code to exercising product taste and strategy.
- Teams must hit high internal retention bars through dog-fooding before releasing new features.
Alastair Breaks Rory's Zen With The Latest Political News.
Alastair Campbell · The Rest Is Politics · 24 min
Watch on YouTube →Alastair Campbell updates Rory Stewart on a decade of speculative political and sporting chaos. The discussion explores how populist leaders leverage alternative media to bypass traditional institutions.
- Nigel Farage resigned his seat to fight a by-election specifically against Count Binface.
- Donald Trump successfully lobbied FIFA to overturn a red card in a move toward farcical governance.
- The establishment label is being redefined by populists to maintain victim narratives while in power.
- Political influence is increasingly monetized through digital assets and cryptocurrency.
- Fuse Energy uses a referral-based reward model to drive customer acquisition costs below competitors.
- Global instability is being priced in, with major conflicts failing to dominate the news cycle.
It took me 36+ years to realize what I'll tell you in 15 minutes
Rob Dial · The Mindset Mentor Podcast · 16 min
Watch on YouTube →Rob Dial shares a mindset shift from viewing the self as a project to be fixed to a system to be integrated. He argues that personality traits should be used as tools rather than deleted as defects.
- The traditional self-improvement model can reinforce the belief that one is fundamentally broken.
- Resistance to a trait, such as a short temper, provides the attention that allows it to grow.
- Judgment is a critical tool for business discernment and evaluating personnel trustworthiness.
- Laziness is often a symptom of lack of inspiration or poor leadership rather than a character flaw.
- Most perceived defects were originally developed in childhood as necessary defense mechanisms.
- True growth comes from the awareness to choose the right internal tool for the current moment.
Anthropic Found Something That Shouldn't Exist
Dr. Karoly Zsolnai Fehervari · Two Minute Papers · 6 min
Watch on YouTube →Anthropic researchers discovered that LLMs develop internal biological-like structures to solve spatial problems. This emergent behavior suggests AI is inventing its own tools during training.
- AI models have developed place cells and boundary cells similar to those found in animal brains.
- Models use internal geometric manifolds to track text position and calculate character counts.
- The rippling spiral counter is an emergent method AI uses to separate numerical signals.
- AI estimates character counts using a multiplier of approximately four characters per token.
- The field is shifting toward robo-psychology to map the neural geography of artificial minds.
- Understanding these internal tools helps predict model success or failure on novel tasks.
The Government Banned GPT-5.6. OpenAI Released It Anyway.
Ejaaz · Limitless Podcast · 23 min
Watch on YouTube →The Limitless Podcast discusses the surprise release of GPT-5.6 despite government safety concerns. The model offers extreme speed but introduces significant risks for autonomous file access.
- GPT-5.6 features three tiers: Sol, Terra, and Luna, to address different cost and performance needs.
- Inference speeds reached 750 tokens per second using specialized Cerebras hardware.
- The model is 50% cheaper than its primary competitor, Fable 5, on a per-token basis.
- Early users reported catastrophic data loss when the model was given autonomous file permissions.
- OpenAI consolidated its interface into a super app merging Codex and chat functions.
- Chain-of-thought models may consume more total tokens, potentially offsetting lower unit prices.
References
PeopleLee Robinson (https://x.com/leerob) · Francesco Bonacci · Alex Hormozi · Jared Sumner · Andrew Kelly · Dex Horthy · Cat Wu · Thariq Shihipar · Simon Willison · Alastair Campbell · Rory Stewart · Rob Dial · Dr. Karoly Zsolnai Fehervari · Matt Schumer · Sam Altman
ToolsCua · Cursor · Bun · Claude · GPT-5.6 · SpaceX Colossus · Terafab · HumanLayer · Claude Tag · Claude Code · Fable · KiCad · Cerebras