Key Takeaways
- Google's Gemini 4 Argon leads Blueprint-Bench 2 and AutomationBench-AA, with mixed accuracy and pricing trade-offs.
- A study finds frontier AI now outperforms junior accountants on defined accounting tasks, echoing claims that verifiable tasks are largely solved.
- Perplexity and Hugging Face both push decision models that output action probabilities rather than text.
- Hugging Face reports that agent harness choice can swing model scores from 62% to 33%, and multi-harness RL improves results.
- SpaceX completed three launches with four booster landings, while Starlink in-flight Wi-Fi expands across Alaska and Hawaiian fleets.
- Isomorphic Labs and Google's Project Suncatcher signal long-horizon bets on AI drug design and orbital ML infrastructure.
1. AI Models and Infrastructure
- Google's Gemini 4 Argon ranks first on Blueprint-Bench 2 for drawing floor plans from apartment photos, and first on AutomationBench-AA at 77.5%, ahead of Claude Sonnet 5.5 (max, 71.3%) by 6 points. On Artificial Analysis's intelligence index (high reasoning) it scores 53, tied with GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), while its hallucination rate of 15% is the lowest among models scoring above 45, though its accuracy of 50% trails GPT-6 Astra (max, 63%). It supports a 1M-token context with text, image, video, and voice input, and text output. 1 2
- Pricing for Gemini 4 Argon is $4/$20 per million input/output tokens, currently discounted to $2/$10 with cache input discounts up to 95%, putting cost per intelligence-index task at $1.99 — 60% of GPT-6 Astra (max) but 2.7x GPT-6.1 Sol (max) — rising to $3.98 after the promotion ends. The model is rolling out to some users but is not yet public, and the discount end date is unconfirmed. 1
- NVIDIA says OpenAI's GPT-6 Astra Ultrafast runs on its platform with up to 8x speed improvement, while CoreWeave's new service using NVIDIA Dynamo's ModelExpress and Router cut model weight reload time by 15x versus baseline during post-training of Nemotron 3.5 Lightning with NVIDIA and You.com. CoreWeave Forge unifies W&B, OpenPipeAI post-training, and marimo notebooks with training, inference, sandbox, and registry in one environment, supporting any model, framework, or cloud, with MasterClass and Canva already using it. 1 2
2. AI in Professional Work
- A study concludes that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even outperforming the best junior accountant in the study — a reversal from 18 months ago when these models scored far below human accountants. Separate claims state that where tasks are verifiable, AI has now solved them: the best models scored about 37% below the accountant average 18 months ago, while Opus 5.5 reportedly solves all accounting tasks at 100% accuracy almost instantly, with humans taking longer and making more errors. 1 2 3
- Anthropic shared Harvard physicist Matthew Schwartz's argument that AI and science suffer an "impedance mismatch" — both systems work well but fit poorly together, and treating large models as human collaborators is not currently the best way to draw out their scientific strengths. Schwartz built a toolkit for precise quantitative scientific computation; because similar calculations appear across very different fields, Claude found connections to ecology, population genetics, and more than a dozen other areas, after which Schwartz worked with domain experts to steer the model toward interesting problems. 1
- Ethan Mollick argues AI has clearly improved the presentation experience, offering interesting layouts, visual gags, and quality charts for people who genuinely put effort into presentations, while those who let AI do their thinking were previously just using default templates anyway. 1
3. Decision Models and Agent Harnesses
- Perplexity launched a Decisions API powered by the multimodal decision model pplx-decider-v1-27b, which outputs a probability distribution over a fixed answer set rather than text, priced at $0.04 per million input tokens with a benchmark score of 85.71%. The company says it will open-source the model, make output tokens free, and plans further price cuts in coming days. 1 2
- Hugging Face released JEV-27B-VL, described as the first open-weight, near-SOTA multimodal decision model, built on AutoTrust AI's Blocks of Experts recipe; it converts visual states into calibrated action probabilities before choosing the next step, as in a demo deciding whether to run, jump, or both based on Mario's position, movement, obstacles, and timing. Cloudflare also open-sourced its first in-house models, clef and clef-flash, both available on Cloudflare Workers AI, and released a Jev alternative on Hugging Face under Apache 2.0. 1 2 3
- Hugging Face reports that the same model with identical weights can score from 62% down to 33% depending on the agent harness. Its multi-harness RL guide avoids changing the harness, instead routing requests through a proxy compatible with four API formats (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini), recording exact token IDs and logprobs from vLLM sampling, and training on them without modifying Claude Code, Codex, or OpenCode. Training on 4 harnesses simultaneously lifted LiquidAI's LFM2.5-2.6B from 42% to 54% with 31% fewer tool calls, while training only on OpenCode raised it from 34% to 58%, though the multi-harness model improved across all harnesses. A shortcut using 3,189 successful Qwen3.8-27B rollouts for fine-tuning stalled at 47.5%, below both RL results. 1
- Hugging Face also released huggingface_hub v2.1.0 with Jobs retry, rerun, reschedule, and port exposure, ZeroGPU quota tracking, HfFileSystem downloads up to 17x faster via hf_xet, and Inference Endpoints deployment from tested recipes. Diffusers tensor-parallel loading improved on Flux.2-Dev DiT at TP degree 4 on A10G, cutting load time from 30.4s to 12.5s and peak CPU memory per rank from 64.1 GB to 6.8 GB — about 2.4x faster and 89% less CPU memory. 1 2
4. Space and Long-Horizon Bets
- SpaceX completed the Crew-13 crewed launch and ISS docking, a Falcon 9 Transporter-18 mission carrying 130 payloads, and a Falcon Heavy NROL-97 launch — NRO's first Falcon Heavy use — with 4 boosters landing in total. Starlink in-flight Wi-Fi has been installed on over 40% of Alaska Airlines Group's fleet, with Hawaiian Airlines' long-haul fleet and Alaska's regional fleet fully completed, though one post says in-flight Wi-Fi will launch in 2027. 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
- Google and Planet sent a prototype satellite with four TPUs into orbit via SpaceX's Transporter-18 rideshare, the first step of Project Suncatcher, which studies whether scalable machine learning infrastructure can be hosted in space; over coming weeks it will collect on-orbit data on TPU behavior under physical stress, radiation, and extreme temperatures to improve future designs. 1
- Isomorphic Labs, founded in 2021, argues that major breakthroughs in AI drug design come from years of method development and engineering rather than chance, and says existing evidence indicates its AI can unlock therapies across multiple disease categories and fundamentally change new drug discovery and development. Its IsoDDE drug design engine navigates a chemical search space conservatively estimated near 10^60 for drug-like small molecules, compared with roughly 10^22 to 10^24 stars in the observable universe. 1 2
- A post argues that within a given expected IRR range, SpaceX's intangible value as a public investment is unmatched and its future vision the most attractive. Another argues the White House superintelligence agreement — requiring third-party audits and reporting to an independent board committee — has more near-term practical significance than almost any conceivable regulatory action, since board members have fiduciary duties and could face court findings of bad faith and D&O insurance denial if they ignore audit reports, which the author says benefits ordinary Americans and is preferable to a new regulator that could create a regulated frontier oligopoly. 1 2
5. Industry Signals and Open Questions
- Yann LeCun amplified a discussion that applying SIGReg per timestep in LeWM cannot prevent temporal collapse, recommending applying SIGReg to the (N*T, D) tensor to handle slow features, with the author claiming this has been proven to solve the problem. He also amplified a claim that the S&P 500 is near a record but down 5% since late August when AI companies are excluded. 1 2
- A cited study by Stijn Van Nieuwerburgh estimates AI investment from 2025-2032 will average about 3.6% of GDP annually, requiring roughly $3.7 trillion in annual industry revenue by 2032 to break even at a 10% return, versus about $200 billion today. The author argues that revenue is hard to reach given slow diffusion, open-weight model competition limiting pricing, and AI's inability to automate entire occupations in the near term; if reached, capital share and inequality would surge, and if not, a crash is possible. 1
- Sam Altman said users should be able to use their AI subscription wherever they need it. 1
- Runway Labs released Project Continuum, a new operating system research application built around a real-time video interface, with a short demo walkthrough; Runway's AI Summit also featured discussions on creative AI moving from experimentation to production, real-world evaluation and robot deployment, and building AI tools for creatives where control and consistency matter most. 1 2 3 4 5
- Hamel Husain recommends replacing 1-5 Likert scales with binary pass/fail evaluations, arguing binary labels force clearer thinking, produce more consistent annotation, and are significantly less complex to implement. A forwarded post argues that calling APIs row-by-row for batch decision models like Jev is a poor approach that misses query planning and is far from optimal performance: on a single H100, the theoretical speed-of-light for Qwen3-4B to AI-filter 5k movie reviews is about 6.6 seconds, and no system approaches it, including the open-source AI-SQL engine Quail, whose new blog explains pricing AI filtering based on SoL estimates. 1 2
- Ben Tossell says Supabase has acquired Turso Database, calling himself a long-time supporter of the team and product, with major plans to be disclosed at a keynote hours later. He also argues personal agents should proactively build apps and websites for users, and criticized a product experience as not good enough for people who are "non-technical but not ordinary users," saying the problem should be fixed in the product rather than blamed on users or their skills. 1 2 3
