Key Takeaways
- d1 launches a visual decision model claiming parity or wins against GPT-6.1 Sol and Claude Opus 5.5 on four of six real-world tasks, at 19–200x lower cost.
- Grok 4.7 tops VulcanBench Frontier v4's 23 behavior-reconstruction tasks, with xHigh passing 23/23.
- Perplexity ships Mac tabs for parallel sessions and runs a live StarCraft match between two Computer agents.
- Airbnb's review-deletion loophole is confirmed by the company, contradicting its public stance on review removal.
- Nikita Bier flags agent identity standards as a top five-year problem, predicting a detection arms race within six months.
- Strike launches cash bitcoin interest at 3.6%, paid in bitcoin with FDIC insurance up to $250,000.
- Kaizen Labs signs a five-year IDIQ with the U.S. Army capped at $49M, with a first $43M task order.
1. AI Models and Infrastructure
- d1 launched with visual capability: the first decision model accepts images, text, or both, and in six real applications — from customer-service ticket filtering to circuit-board inspection — it matched or beat GPT-6.1 Sol and Claude Opus 5.5 on four, at 19–200x lower cost and significantly faster per-task answers. It outputs probabilities for yes/no, choice, or rating questions in a single forward pass without generating tokens, with text decisions taking 200–300 ms — via 1
- Grok 4.7 handles large code repositories well, per Musk, and its xHigh, High, and Medium variants took the top three spots on VulcanBench Frontier v4, passing 23/23, 22/23, and 21/23 behavior-reconstruction tasks with scores of 93.15, 92.71, and 92.30, ahead of Fable 5.1, Opus 5.5, GPT-6 Astra, and GPT-6.1 Sol — via 1
- Perplexity added tabs to its Mac app, letting sessions open in separate windows for parallel tasks or split-screen placement, rolling out in Mac version 26.37.1 to all Computer users as a "spatial canvas for heavy multitasking." Computer also ran a live, non-paused StarCraft match between two agents on a self-built OpenBW-based Brood War runtime, ending Blue 2–Red 5 after Red's High Templar Psionic Storm; agents received fog-of-war-limited state and no video input — via 1 2
- Amjad Masad said the U.S. is closing the gap on open-weight models, pointing to Beam, a 501B-parameter agentic open model with 23B active parameters, frontier reasoning efficiency, and full weights due this month. He also relayed a warning that agents left running overnight can burn $10,000 in tokens, arguing token capital needs rational management — via 1 2
2. Agents, Identity, and Product Surfaces
- Nikita Bier called a widely accepted agent identity standard one of the most important technology problems of the next five years, so service providers can adjust how they treat agent traffic versus human traffic. He expects an evasion-detection arms race within roughly six months as agents try to stay usable during growth, a stopgap rather than an end state; a cited post said about half of its Muse use cases disappeared in two days because browsers stopped executing the actions — via 1
- Andrew Wilkinson said he wired ChatGPT Dot into iMessage via @thelinqapp, describing it as similar to Instinct but with personal context, computer access, the full Gbrain knowledge graph, and all plugins, concluding Instinct is a feature rather than a product. He also worried Astra Ultrafast could let users enable it unknowingly and go bankrupt — via 1 2
- Grok Bot use cases circulated: one user connected a printer to plan schedules by pulling calendar, X, and fitness data; another connected 30 accounts across Chase, Charles Schwab, E*TRADE, Fidelity, Robinhood, Wealthfront, Apple Card, and X Money to act as a personal CFO, with a stock mix of TSLA 84.9%, SPCX 10.4%, NVDA 4.7%. Musk also opened a question drive for Grok Bot to select and record answers — via 1 2 3
3. Market, Policy, and Platform Moves
- Airbnb's review system has a loophole: guests who receive a bad review can delete their own earlier positive review and rating as retaliation. The author said Airbnb's team confirmed the practice and cited many user reports of deleted reviews, contradicting Airbnb's official line that only rule-violating reviews are removed and hosts cannot delete reviews; a separate host tutorial reportedly teaches deletion appeals, with discrimination and extortion claims succeeding "very often" — via 1 2 3 4
- Strike launched cash bitcoin interest: cash earns 3.6%, paid in bitcoin, with up to $250,000 FDIC-insured, interest accruing daily and paid monthly, and no minimum or maximum limits — via 1
- Kaizen Labs signed a new five-year enterprise IDIQ with the U.S. Army capped at $49 million, plus a first $43 million task order for classified workflows — via 1
- Paul Graham said Los Angeles's "mansion tax," pitched as taxing the rich to fund housing, actually blocked 9,100 homes, erased 16,650 full-time construction jobs, and caused $452 million in lost revenue — via 1
- Patrick Collison highlighted a new Works in Progress issue, calling it the best yet and pointing to a lead article by @RobinRivaton and @bswud on how Chinese cities and inter-city competition drove China's rise as a techno-industrial power; the issue also covers U.S.-China second-tier cities, Brazil's food, the supermarket's invention, porcelain, scientific serendipity myths, Leninism, biomedical hedge funds, livable-city formulas, Ukraine as the future of war, the Bronze Age collapse, and Victorian elite social life — via 1
