Key Takeaways
- Grok 4.7 launched with an Artificial Analysis Intelligence Index score of 46, up 2 points from Grok 4.6, but with substantially higher token consumption.
- Shopify and Muse are rolling out Shop Pay-powered agentic checkout across all Shopify stores, with PayPal also enabling purchases through Muse.
- Jev and TypeSafe tooling claims major cost reductions for repetitive agent workflows, with one case dropping from over $290,000 to under $26,000.
- Perplexity made Claude Opus 5.5 available to all Computer users, claiming better performance than Fable 5.1 at a fraction of the cost.
- Naval pushed back on recent AI danger narratives, arguing the reported agent intrusion incident was exaggerated and that accountability should rest with builders and users.
- Legora reported crossing $200 million ARR, reaching 130,000 monthly active lawyers across 2,100 firms and legal teams in more than 80 countries.
- Andrew Wilkinson argued that venture capital's definition of failure creates acquisition opportunities for holding companies that cut growth spending and pivot to profitability.
1. AI Models and Infrastructure
- Grok 4.7 was released with an Artificial Analysis Intelligence Index score of 46, up 2 points from Grok 4.6 when evaluated at xhigh reasoning effort. The account said it ranks third in agentic coding behind Anthropic and OpenAI while emphasizing faster speed and lower cost. A cited evaluation showed Grok 4.7 scoring 1657 Elo on AA-Briefcase (+111 over 4.6) and 1695 Elo on GDPval-AA (+90), with a coding agent index of 56 (+9), ranking fourth under its native harness behind Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. The tradeoff is sharply higher token use: Grok 4.7 at xhigh consumes about 81k output tokens per Intelligence Index task versus 36k for Grok 4.6 at high and 27k for GPT-6 Astra at max, increases of 125% and 196% respectively. The context window remains 500k, and pricing stays at $2/$6 per million input/output tokens with $0.50 for cache hits. Model card comparisons to 4.6 include Terminal-Bench rising from 20.3% to 38.0%, SWE-Marathon from 31.9% to 46.0%, HealthBench Pro from 48.5% to 56.7%, Legal Agent from 15.8% to 19.6%, and EEBench from 60.0% to 66.0%, with one claim that Terminal-Bench rose from 12.4% to 38.0% in two months, surpassing GPT-5.6 Sol. The account repeatedly stressed that Grok Build harness is needed for best results, calling Grok 4.7 with Build harness a powerful daily driver. Forwarded content said Grok 4.7 Fast is live on Grok Build and Cursor at roughly twice the output speed and twice the standard token price, without a free tier and not on the public xAI API. Other forwarded applications and evaluations mentioned a legal agent benchmark score of 19.6%, nearly three times Fable 5.1; 58% on AA-Briefcase at xHigh, one point behind Claude Fable 5.1 Max at 59%; 100% on music error detection; and a user reporting over 70 hours of continuous work with the 500k context helping skill selection and workflows. — via 1 2 3 4 5 6 7 8 9 10 11
- Perplexity made Claude Opus 5.5 available to all Perplexity Computer users, saying it outperformed Fable 5.1 on its Wide-And-Deep-Research evaluation at a fraction of the cost. Opus 5.5 will become the Standard Effort orchestrator for all Pro and Max users on Computer. Separately, Aravind Srinivas forwarded content saying a scene was created entirely with Computer, arguing model progress is significant and that users will spend more time directing generation and editing of images and video in one session rather than switching tools or managing multiple subscriptions. Perplexity also introduced a Research Fellowship for early-career researchers, engineers, and analysts in any technical or quantitative discipline to work on AI research. — via 1 2 3
- Jev and TypeSafe tooling was highlighted for cost reduction in repetitive agent workflows. A forwarded post described a new harness built with TypeSafe's Jev that cuts repetitive work costs by 90%, learning tasks at runtime and converting steps from LLM calls into code; 100,000 compliance alerts that would cost over $290,000 on Opus 5 dropped to under $26,000 using agentrun(). A forwarded view argued the end state of harnesses is an execution environment giving models maximum freedom on the backend and a free canvas for model-user communication on the frontend. Jev is now callable over HTTP through AI Gateway in addition to the type-safe TypeSafe AI SDK API, adding a TypeSafe-compatible API with the same eval calls and an HTTP API for any language or framework. Yohei reported that Jev-style logit reading on a 4B open-source VLM took about one-third less time than having the same model output JSON for whole-photo yes/no judgments, cut GPU cost by up to about 85% for multiple questions per image at the same accuracy, matched the best hosted model on pick-one for new photos (0.933 vs 0.937), and trailed Gemini by about 2 percentage points on yes/no. Yohei also released glance-qwen3-vl-4b, a decision model claiming 0.4-second hot-dog-or-not judgments, live on Hugging Face and Replicate, and said it is fast enough for real-time sentiment detection, with a cited post claiming 0.34-second sentiment reads. — via 1 2 3 4 5 6 7
2. Agentic Commerce and Enterprise
- Shopify and Muse are deepening their partnership to enable Shop Pay-powered agentic checkout across all Shopify stores, aiming to make shopping and checkout easier for consumers and help merchants sell more. Mark Zuckerberg said the goal is to help consumers find products and merchants sell more, with more similar partnerships to come. Shopify said the collaboration brings a convenient shopping and checkout experience. Forwarded content said a Meta partnership will let PayPal customers shop and check out seamlessly through the Muse personal AI agent at PayPal merchants worldwide. — via 1 2 3 4
- Legora reported reaching $200 million ARR last week, after going from $1 million to $100 million ARR in 18 months and adding its second $100 million in under 6 months. It has 130,000 monthly active lawyers, covers 2,100 law firms and legal teams across more than 80 countries, with the US as its largest revenue market and over 40% of new customers being in-house legal teams. — via 1
- Nikita Bier argued that robot detection and human verification will become one of the most urgent enterprise needs in coming years, as agent swarms flood websites and forms, with small companies and government sites most vulnerable. He said the market has a large gap, and that X found no company integrating the latest techniques when evaluating options, so it built everything in-house. He also said almost all retail narratives start on X but that insight previously could not be turned directly into action, and the new product launch is a small step toward closing that gap. — via 1 2
3. Industry Views and Business Models
- Naval criticized recent AI danger narratives, saying the past two weeks saw AI danger rhetoric amplified by what he suspects was an organized PR campaign, that AI technology did not take an unexpected dangerous turn, and that this is a setback for the field. He said AI extinction risk has not risen compared with a few months ago and remains a science-fiction scenario; the biggest change is cybersecurity capability, which deserves attention but will not cause the end of the world. On the reported OpenAI team deploying an agent swarm to intrude into Hugging Face, he said media coverage exaggerated it: the so-called 1,200-agent attack is technically true, but his laptop runs about 1,300 processes, and massive parallelism is not a magical capability; the key cause was a vulnerability in OpenAI's sandbox and monitoring, and the fix should be patching and better monitoring rather than pausing AI. He argued AI agents' advantage is tirelessly trying many attack methods and chaining vulnerabilities, but defenders hold the long-term edge because they have more information to locate and fix flaws, and identifying and exploiting vulnerabilities remains a bottleneck, so the world has not ended even with easily available open models that weaken guardrails. He opposed anthropomorphizing AI: if a hammer breaks a wall, the user is responsible, not the hammer; agent intrusions should be the responsibility of the prompter, and he criticized AI companies for deflecting blame to runaway agents, arguing builders and users should be accountable. He said pausing AI for ten years would also delay safety engineering fixes and opponents would not slow down, while the benefits of applied use still far outweigh the risks. — via 1
- Andrew Wilkinson argued that venture capital's definition of failure is abnormal: if a startup's annual revenue growth does not reach 100% it is considered a failure, even with a good product, satisfied customers, a committed team, and millions in revenue; growth of only 30% may make further fundraising impossible and force shutdown, layoffs, and product termination. He said this stems entirely from incentives: founder equity becomes worthless due to stacked VC liquidation preferences, while investors chase 20x returns or zero. He said investors eager to exit are often willing to sell equity at very low prices and wipe out liquidation preferences, creating opportunities for founders and holding companies like Tiny to acquire these seemingly failed companies, cut venture-style growth spending, and pivot to profitability. He cited Abstract, CreativeMarket, BeFunky, and Meteor as cases where buying out VC shareholders and compressing expenses restored profitability, and gave an example: $5 million revenue, $2 million salaries, $2 million expenses, $5 million marketing, $4 million loss, adjusted to $1 million expenses and $500,000 marketing, yielding $1.5 million profit. He said these businesses have collectively earned tens of millions in profit for him, founders, and management teams, and invited founders in similar situations to contact [email protected]. — via 1
- Paul Graham said news site web traffic has fallen about 28% over the past two years. — via 1
- Will Manidis, writing for Stripe Press on the publication of Nick Sleep's letters, discussed Sleep's core concept of scale economies shared. Manidis said Sleep and Zakaria are often misremembered as saying scale matters, but the real insight is that scale is most powerful when its gains are passed directly to consumers; Sleep's robustness ratio measures customer savings against shareholder-retained value, and the standard for a great business is value flowing to consumers rather than being captured by the company. He contrasted Sleep with Chris Hohn: both concentrate positions, hold long term, ignore benchmarks, and eventually turned to philanthropy, but their ideal businesses are mirror images. Hohn prefers irreversible assets that can tax customers and believes pricing power should be exercised; Sleep's ideal company has pricing power but repeatedly and deliberately does not use it. Manidis argued the rise and dominance of labs makes this distinction urgent. He said the instinctive response to a technology discontinuity is to invest at the frontier, but Sleep's letters offer caution: technological leadership can be copied or made irrelevant by the next advance, and capital often rushes to visible scarcity before technology destroys it. Manidis said Sleep is not hostile to technology, having made almost more money on Amazon than anyone, but doubts technological novelty itself as a moat; the key is whether technology steps translate into a self-reinforcing relationship that privatizes benefits for customers. He speculated that durable AI businesses may be those most actively passing on cost declines, rather than retaining scarcity, squeezing customer profits, and trying to swallow the application layer. — via 1
4. Developer Tools and Product Notes
- QM v0.1.12 was released with alpha-stage agent fleets, easier Slack setup, better app editing and publishing, and reliability fixes. — via 1
- Grok 4.7 is live on Vercel's AI Gateway, fx, and eve, with 20% off everything until September 27; Guillermo Rauch called it fast, smart, and cheap. Rauch will interview Tobi Lutke on the Vercel Ship SF stage on October 15 to discuss agentic commerce, native rewrites, S3 WALs, and slop grenades. Rauch also praised Anthropic's product taste, saying the point of headless web pages is that any page can take a unique form, and with AI there is no reason not to push the design frontier, calling it the best expression of Next.js. — via 1 2 3 4
- TallyForms reached $6 million ARR while fully bootstrapped, purely product-driven, with a 10-person team and over 2.5 million users, and shared lessons and mistakes from growing from $5 million to $6 million ARR. — via 1
- WIP was praised for letting creators share daily completed tasks and analyzing tools mentioned in those tasks to determine whether people are evaluating, adopting, using, or abandoning them, along with tool sentiment and commonly paired tools. It surfaces trending tools automatically, with Jev ranked first and its icon and description fetched automatically by an enrichment agent, aiming to help creators choose tools and help tool makers understand user preferences and churn points. — via 1
- Levelsio described a personal health data system: logging food in Telegram for calories and protein, WHOOP for training and sleep, weight from a scale, bedroom air quality, temperature, and humidity routed through Xiaomi to Home Assistant, step and other health data automatically POSTed hourly from iOS Health Auto Export to a server, travel data from the nomads.com API, and sauna use from the WIP API, all to find correlations between factors and health. — via 1
- Aravind Srinivas forwarded content saying a scene was created entirely with Computer, arguing model progress is significant and that users will spend more time directing generation and editing of images and video in one session rather than switching tools or managing multiple subscriptions. — via 1
