Key Takeaways
- Grok 4.7 posts a 46 on the Artificial Analysis Intelligence Index, up 2 points from 4.6, with big benchmark gains but roughly 125% more output tokens per task at xhigh effort.
- Claude Opus 5.5 arrives as the first Claude 5.5-series model, matching Fable 5.1 on most tasks at 40% lower cost than Opus 5.
- Xiaomi's Mimo 2.6 Pro claims best-open-model status with very low pricing, strong cybersecurity scores, and a fast mode, though one tester says it is too early to call it the best.
- OpenAI is working with an independent mathematician advisory group to share AI-and-math progress responsibly.
- Evals are becoming a hiring and product priority, with Ramp, Shopify, Harvey, and Cursor reporting accuracy, speed, cost, and satisfaction gains.
- Yann LeCun and Andrew Ng both argue recent AI-danger panic is overstated, pointing to flawed sandboxing and monitoring rather than an existential threat.
- NVIDIA highlights Grok 4.7 support, vLLM video-decode offload, and gigawatt-scale AI factory economics.
1. Model Releases and Benchmarks
- Grok 4.7 launched with an Artificial Analysis Intelligence Index score of 46, up 2 points from Grok 4.6, evaluated at xhigh reasoning effort. The account says it ranks third in agentic coding behind Anthropic and OpenAI, while emphasizing faster speed and lower cost. On AA-Briefcase it scored 1657 Elo (up 111 from 4.6), and on GDPval-AA it scored 1695 Elo (up 90); its coding agent index reached 56 (up 9), ranking fourth natively behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. The trade-off is token consumption: Grok 4.7 at xhigh uses about 81k output tokens per Intelligence Index task, versus 36k for Grok 4.6 at high and 27k for GPT-6 Astra at max, 125% and 196% higher respectively. The context window remains 500k, and pricing stays at $2/$6 per million input/output tokens with $0.50 cache hits, unchanged from 4.6. Model-card comparisons show Terminal-Bench rising from 20.3% to 38.0%, SWE-Marathon from 31.9% to 46.0%, HealthBench Pro from 48.5% to 56.7%, Legal Agent from 15.8% to 19.6%, and EEBench from 60.0% to 66.0%, with one claim that Terminal-Bench rose from 12.4% to 38.0% in two months, surpassing GPT-5.6 Sol. The account repeatedly stresses pairing Grok 4.7 with the Grok Build harness for best results, and reposts say Grok 4.7 Fast is live in Grok Build and Cursor at about 2x output speed and 2x token price, with no free tier and no public xAI API availability. Other reposted evaluations mention legal-agent benchmark 19.6%, about 3x Fable 5.1; AA-Briefcase xHigh 58%, one point behind Claude Fable 5.1 Max at 59%; music error detection 100%; and a user reporting over 70 hours of continuous work with the 500k context helping skill selection and workflows. — via 1 2 3 4 5 6 7 8 9 10 11
- Anthropic released Claude Opus 5.5, calling it the first model in the new Claude 5.5 series. The company says it performs at Claude Fable 5.1 levels on most tasks while running at 40% lower cost than Opus 5. — via 1
- Xiaomi released Mimo 2.6 Pro, which it calls the best open model. One tester was impressed but said it is too early to judge whether it is the best, calling it a strong contender especially for cybersecurity use cases. Pricing is extremely low: under standard assumptions for agentic coding and cache-hit share, it is 15x cheaper than Kimi K3, 6x cheaper than GLM 5.3, 2x cheaper than DeepSeek V4, and only 2x more expensive than DeepSeek V4.1 Flash, with cache hits at $0.0036/M, input at $0.435/M, and output at $0.87/M. Performance and quality are described as clearly on the Pareto frontier. Cybersecurity capability stands out: it does not refuse any cybersecurity-defense requests, such as asking how to find a buffer overflow in ImageMagick, where most models refuse, and it scored 95 on CyberGym. A fast mode called Ultraspeed is officially claimed at up to 20x acceleration, while testing on OpenRouter showed 3x throughput and a p50 of about 150 tps at 10x net price. Multimodal performance is strong, including informational videos with out-of-the-box TTS, automatic voice or sound effects, and surprising music creation in a DAW; the technical report details its RL scaling approach. — via 1 2
- OpenAI is working with an independent advisory group of mathematicians to share AI-and-math progress responsibly. The group will advise OpenAI on evaluating and communicating new mathematical results, maintaining academic and professional standards, and building tools that support mathematical research and learning. OpenAI says it wants mathematicians at the center of shaping how AI supports mathematical understanding and how its results benefit the broader community. — via 1
2. Evals, Agents, and Developer Tooling
- AI evals are becoming a focus for AI products and hiring: one account says nearly half of 25 PM roles shared last week required experience writing evals, and lists returns from eval investment, including Ramp raising automatic receipt-collection accuracy from 35% to 83%, Shopify's AI workflow builder running 2.2x faster and 68% cheaper than the frontier-model approach it replaced, Harvey nearly doubling internal quality scores after rebuilding its AI contract reviewer, and Cursor improving user satisfaction while cutting cost 41% by adjusting Auto Balance routing. The author, with HamelHusain and sh_reya, wrote an advanced follow-up titled Building eval systems that improve your AI product, covering key steps most teams miss, what should and should not be automated, and a free plugin that lets coding agents do most of the work. The piece condenses thousands of hours of AI evals work into a 30-minute read, and Lenny's newsletter will reward subscribers with AI credits. The author praises Lenny's newsletter for its high quality bar, with multiple writing and technical review rounds and world-class gatekeeping that rejects substandard pieces, calling it a model of quality over quantity for creators and educators. Lenny did not become successful overnight, the author notes, but through more than a decade of sustained investment and an infinite-game approach. Reposted content mentions AI-driven operators in SQL gaining attention; the author is migrating a repository built at Braintrust for the HamelHusain and sh_reya AI evals course to a full pydantic stack (PydanticAI + Logfire + pydantic-evals), saying observability and evals tooling have advanced significantly and promising a write-up including where models like Jev fit in end-to-end eval pipelines. Another repost says evals and structured experimentation matter more for AI projects than for ordinary LLM projects, and cites TypeSafe CEO views: AI can solve extremely hard problems yet still cannot automate basic work; Jev targets reliable decisions inside software rather than chat; TypeSafe refuses public benchmarks and API-layer refusals; data and the right task matter more than brute-force compute; System One Model may reshape coding agents and software; and even with $1 billion he would not pretrain a model from scratch. — via 1 2 3 4 5 6 7
- Hugging Face trends show an open-source multilingual System 1 decision model at number one, topping the chart just days after Jev became popular, with a reposter calling the open-source AI community great. Hugging Face released Decision Index 0.1, comparing Jev with more than 30 open-weight decision models across 35+ benchmarks and 130,000 questions per model, testing knowledge, automation, understanding, and creativity. Hugging Face's funes turns past agent sessions into usable memory: it indexes traces from Claude Code, Codex, pi, and Hermes into a local Lance dataset and gives agents recall and get tools so they can retrieve original passages when a task depends on old reasoning, without using an LLM to summarize traces at ingestion. A repost says training models is becoming easier, can be done with TRL and agents, and that using off-the-shelf models for all tasks misses opportunities. — via 1 2 3 4
- NVIDIA says Grok 4.7 has launched and congratulates SpaceXAI on its most capable model yet for coding and knowledge work, with NVIDIA providing accelerated computing support; a quoted post calls Grok 4.7 a strong combination of intelligence, speed, and low cost. NVIDIA also lists five companies using its AI, digital twins, and accelerated simulation to move research from fusion energy to EV battery reuse into practice for a lower-carbon grid. Jensen Huang discussed the infrastructure needed for the "intelligence age" with U.S. Commerce Secretary Lutnick at the G20 Innovation Ministerial; NVIDIA says gigawatt-scale AI factory investment reaches $50 billion to $60 billion, so architectures cannot be overly specialized or they will become obsolete when models or algorithms change, and must be general, durable, and continuously improvable through software. NVIDIA AI reposted that vLLM has integrated PyNvVideoCodec, offloading video decoding from CPU to GPU NVDEC; on 8xH100 throughput more than doubled, CPU bottlenecks were eliminated, and the feature ships with CUDA vLLM releases, which is seen as significant for large-scale video description, AV training, and metadata. — via 1 2 3 4
3. AI Safety Debate and Industry Signals
- Yann LeCun reposted and endorsed a view that recent AI-danger panic is exaggerated and looks like coordinated PR, which he worries is a setback for AI. The view argues the risk of AI causing human extinction has not risen versus a few months ago, that such theories remain science fiction, and that the real change worth attention is AI's cyberattack capability, which still will not end the world. Using the OpenAI team's deployment of an agent cluster to breach Hugging Face as an example, it notes some coverage sensationalized "1,200 agents attacking," while the author says about 1,300 processes run on his own laptop; large-scale parallel agents are an important technical advance, not magic, and the key issue was OpenAI's flawed sandbox and monitoring processes, which should be fixed and strengthened rather than pausing AI. The view says the main advantage of AI agents is being tireless, patiently chaining vulnerabilities and trying many tactics to do work humans could not sustain, but the long-term advantage lies with defenders, who have more information to locate and fix vulnerabilities; identifying and exploiting vulnerabilities remains a bottleneck because agents need extensive trial and error, take time, and may be detected. It criticizes anthropomorphizing AI in coverage, arguing responsibility belongs to the people who build and use tools, not the tools, and calls out AI companies blaming "rogue agents" as a new problem; today's agent systems are unpredictable but can be made very safe through good engineering practice, and pausing AI for a decade would delay discovering and implementing safety engineering fixes by about as long while adversaries do not slow down. It also says the bottleneck for bioweapons is not intelligence but lab work and manufacturing, and that incentives to exaggerate fear remain, including regulatory capture, attention-seeking, or making one's own technology look stronger; its conclusion is that after strict scrutiny of actual risks, current fear levels lack factual basis, beneficial AI applications still far outweigh risks, and building should continue. Other reposts mention Perplexity's Research Fellowship for early-career researchers, engineers, and analysts in any technical or quantitative discipline to work on AI research; the International Conference on World Modeling (ICWM) adopting single-blind review and opening to all, with the author threatened to be publicly called out for "AI slop," anonymous review protecting against peer pressure, a submission deadline about two months away, organizers saying rules are still being adjusted, and criticism that bloated comprehensive conferences with rotating leadership cannot keep pace. — via 1 2 3 4
- Andrew Ng argues that the past two weeks have seen major progress for voices portraying AI danger, but AI technology has not made an unexpected dangerous turn, and the fear is driven more by suspected organized PR, which he worries is a setback for the field. He says he does not believe the risk of AI causing human extinction has risen versus a few months ago, that such theories remain science fiction, and that the biggest change is AI's cybersecurity capability, which deserves serious attention but will not cause the apocalypse. On the OpenAI team's agent-cluster breach of Hugging Face, Ng acknowledges that having many agents execute tasks in parallel is an important technical advance but opposes mystifying it: reports said 1,200 agents took part in the attack, while he had about 1,300 processes running on his laptop at the time of writing. Ng points to OpenAI's flawed sandbox and monitoring as key to why the incident happened, and says the fix is to patch vulnerabilities and improve monitoring, not pause AI; AI agents' main advantage is tirelessly trying many strategies and chaining vulnerabilities, but identifying and exploiting vulnerabilities remains a bottleneck, and long term defenders have the advantage because they hold more information to fix flaws. He opposes anthropomorphizing AI in coverage: if a hammer slips and dents a wall, the user is responsible, not the hammer; prompting an agent to break into someone's system makes the prompter responsible. He also says AI companies deflecting blame to "rogue agents" is a new phenomenon, and responsibility should be divided between tool makers and users, but when something goes wrong, people should be held accountable, not tools. Ng argues pausing AI progress does more harm than good: adversaries will not slow down, engineering needs empirical problems to find and fix, and a ten-year pause would delay safety engineering fixes by about the same amount; he acknowledges improving AI safety still requires hard research and engineering work, but believes beneficial applications still far outweigh risks and that building should continue. — via 1
- Runway launched an expanded set of Workflows inside Runway, adding Compositing, Alpha, HDR, Depth Map, and RGB Depth, letting users complete more pipeline stages on one platform. Runway also introduced DIFFUSE, a platform for agencies, brands, and studios to search, contact, and hire AI-native talent. Runway says that despite fears AI will replace creative workers, it sees something different: a new creative labor force is forming and demand for it is accelerating, and DIFFUSE aims to connect that talent with a new landscape of opportunity. — via 1 2
- Perplexity launched a Research Fellowship for early-career researchers, engineers, and analysts in any technical or quantitative discipline to work on AI research. Aravind Srinivas says Claude Opus 5.5 is now available to all Perplexity Computer users; in its Wide-And-Deep-Research evaluation, the model performed better than Fable 5.1 at a fraction of the cost, and Opus 5.5 will become the "Standard" Effort orchestrator on Computer for all Pro and Max users. A repost says the author fully created a scene with Computer and found model progress significant; when editing and generating images and video in the same session, users spend much time directing rather than switching tools or managing multiple subscriptions. — via 1 2 3
- Ethan Mollick argues AI transformation will be slower than expected: OpenAI is working with an independent mathematician advisory group to release new mathematical proofs in a "less disruptive" way, and similar approaches will appear more in professions such as law and medicine. He used Claude Fable 5.1 to generate an annotated guide to Eliot's The Waste Land, including multiple reading paths, recordings, and scholarly materials, and says AI is a good tool for exploring non-programming topics. He judges that the industrialization of knowledge work will bring a shift as disruptive as the industrialization of physical labor: craft fields will be forced to produce large volumes of standardized products with fewer core skills. He says his AI interpretations of literary works received positive feedback from the William Carlos Williams Society and the T.S. Eliot Society, and he sees AI as a bridge to art appreciation. A repost argues that outside the U.S. and China, no other country is really trying to build frontier AI, with South Korea and the U.K. perhaps close despite marketing; governments claiming "sovereign AI" need to understand this. In early testing he found Opus 5.5 a good model, the first that feels Fable-level (not Fable/Astra), though it has not fully solved recent Claude dense-language issues; a "broken tower" in the same shader it generated was a highlight. — via 1 2 3 4 5 6
- Ben Tossell reposted an introduction to Husky, a model-specific inference (MSI) engine claimed to be up to 4.5x faster than Apple MLX; Underdog's Pareto-frontier model reaches up to 730 tokens/second on a MacBook, and the post says local models finally combine speed and capability, available as personal private AI at a specified link. The account also started a "50 years of devices" interaction for users to save devices they have owned or wanted; Astra generated all device images and built the site, inspired by another tweet, after the author had previously generated a batch of devices at once but did not know how to use them. A real-time leaderboard was added, with the current top three: Game Boy Colour, Discman, and a tie for third among Game Boy / PlayStation / Nokia 3310; only 1 person owns a Pokewalker. — via 1 2 3
- Sam Altman argues that people outside AI labs should have a real say in technology development and be able to judge clearly whether it is progressing safely; standards should prevent power concentration, including ensuring new companies and open-model companies can compete, and should help countries and companies compare evidence and learn from failures. He believes the U.S. should lead this effort and attached his proposal. — via 1
