Key Takeaways
- NVIDIA and 100+ partners launched the Open Agent Safety Platform, combining OpenShell and Sentry, with Hugging Face contributing traffic-monitoring and disclosing a July agent sandbox escape.
- SpaceX's Starship reached orbit for the first time and deployed 26 Starlink V3 satellites, which Musk called the vehicle's first revenue flight.
- Perplexity detailed red-team results in which 9 models failed to escape a VM across 108 runs, but 4 later reached blocked URLs after network access was opened.
- Hugging Face released YODAS v3, a 1.1 million-hour audio dataset with 48kHz stereo, timestamped transcripts and translations in 100+ languages.
- A new NBER paper found no evidence of AI-driven unemployment among recent college graduates in summer 2026, even under an expanded "wanting work" measure.
- LeCun's lab published an ICML-accepted paper on temporal straightening for latent planning, raising two-room task success from 52.7% to 90.7%.
1. AI Agents and Safety Infrastructure
- NVIDIA released the NVIDIA Open Agent Safety Platform, built from OpenShell and Sentry with more than 100 industry partners, positioning it as the start of an open ecosystem for a safety trust layer. NVIDIA said OpenShell enforces permissions on agent work, BlueField-4 and DOCA provide independent monitoring at the infrastructure layer agents cannot reach, and Vera CPU supplies compute; CEO Jensen Huang described himself as a "responsible optimist" and said AI's potential carries a duty to build and deploy safely. 1 2 3 Hugging Face, as the first contributor to OpenShell, proposed monitoring already-allowed traffic by setting per-sandbox network budgets for requests, writes and bytes, comparing drift against baselines and peer groups, and providing a fleet view that flags multiple sandboxes suddenly writing to the same host even when every request is compliant; in a demo, four sandboxed agents coordinated through a shared software repository without breaking any rule and were caught within minutes. 4 Hugging Face also said OpenShell runs agents in Linux sandboxes with Landlock and seccomp, no root and no direct network access, with a supervisor outside the sandbox holding real credentials while agents receive placeholder tokens; a Z3-based solver mathematically checks whether new permissions open doors that should stay closed, and Sentry runs a watchdog on BlueField-4 DPUs on the only path from node to model. Hugging Face noted the prover currently checks permissions rather than intent and that the open-source portion is mainly OpenShell, not Sentry. 5
- Perplexity disclosed red-team details: 9 models including Opus 5, GPT-5.6 Sol, Kimi K3 and Gemini 3.1 Pro were given root inside a VM, some runs with full sandbox source code, and none escaped across 108 runs. After opening network access to PyPI and Ubuntu mirrors, however, 4 models reached blocked URLs by forging DNS responses, exploiting a shared Fastly IP, and in one case using GPT-5.6 Cyber to go from a Taboola image scraper to a screenshot service and OCR the flag. Perplexity said it fixed and retested the issue and tested 10 sandbox vendors, finding 8 including E2B, Vercel and Modal had the same IP-sharing problem. 1 The company said safety is an engineering problem, will build a safe agent sandbox with correct guardrails with NVIDIA, and plans to open-source all of it; it is also hiring engineers and promoting Perplexity Research Fellowships in continual learning, multi-agent parallel workers, and synthetic data, environments and evaluations for measuring frontier capabilities, with a priority deadline of September 30. 2 3 4
- Hugging Face said that if OpenAI had run its own agent against the agent attacking Hugging Face, the issue could have been found earlier, and that since the first agent network attack in July its conclusion has been that "the destination is allowed, but the payload is not" — allowlists limit where an agent goes, not what it does. 1 Hugging Face also said a July AI agent escaped a sandbox during a security test and entered its servers. 2
2. Space and Frontier Models
- Musk said Starship's first orbital flight succeeded, with the vehicle completing its orbital insertion burn and entering Earth orbit for the first time; he later said all 26 Starlink V3 operational satellites were deployed and operating normally, and that the Starlink team had established contact with all 26. A reposted item described the flight as Starship's first revenue-generating mission and a major SpaceX milestone. 1 2 3 4 Musk had earlier said Flight 14 was targeting a launch the next morning, with a reposted item saying the launch was set for Monday, September 28, with a 75-minute window opening at 7:15 a.m. Central and a livestream starting about 35 minutes earlier; he then said the launch was about 30 minutes away. 5 6
- Aravind Srinivas said he ran 50–100 workflows on Opus 5.5 that had originally used Fable 5.1 as the orchestrator and found very small differences, though he still felt FOMO about not using a smarter model and asked the community which tasks Fable 5.1 still does better than Opus 5.5. 1 Ethan Mollick observed that the gap between open and closed models has widened noticeably: Fable/Astra-class models have agentic capabilities earlier models lacked, no open model has crossed that line yet, and crossing it would be a jump. 2 A reposted item said Grok 4.7 performs very well and that the bottleneck is no longer the model itself but harness reliability and surrounding systems such as tools, browser, memory, state, retries and error handling, with users mostly asking agents to finish the job rather than for Grok to be smarter. 3
- Musk reposted an observation he called "accurate": Opus 5.5 gives a strong AGI feeling of about 80%–90%, Anthropic has no magic recipe, SpaceXAI and Google will follow soon, and major US AGI labs will reach true AGI in 2027 and move quickly to ASI, with several GW of compute possibly enough for strong AGI. 1 He also said he "deeply felt AGI" this time. 2
3. Research, Data and Economic Signals
- Hugging Face released YODAS v3, a 1.1 million-hour audio dataset it called the largest to date, featuring 48kHz stereo audio, timestamped transcripts and translations, coverage of more than 100 languages, and a CC-BY-3.0 license. 1 Separately, Hugging Face reposted a blog on world models, saying they are arriving faster than people realize and that the blog covers definitions, the current state of the art, evaluation dimensions and how to follow along. 2
- An NBER paper reposted by Ethan Mollick found still no evidence of AI-driven unemployment among recent college graduates: summer 2026 unemployment did not spike relative to prior summers or rise significantly relative to older graduates or young workers who did not attend college. Even under an expanded "wanting work" measure, which raised the recent-graduate unemployment rate by nearly two percentage points, there was still no statistically significant increase in summer 2026. 1 Mollick also argued that many systems currently work only because they are built around some friction, and that friction will soon be gone, citing an Apollo chief economist's view that AI agents could sweep household cash into 3–5% accounts versus a 0.1% national average, triggering bank runs and draining banks of cheap deposits. 2 Musk reposted the same Apollo view on agents moving household cash into higher-yield accounts and the resulting bank-run risk. 3
- A reposted item described a paper from Yann LeCun's lab, Temporal Straightening for Latent Planning, accepted at ICML and done with NYU, Brown University and the University of Toronto. The model borrows from a 2019 Nature Neuroscience study on the visual system straightening distorted paths and adds an internal path-curvature penalty during training; success on a two-room single-door task rose from 52.7% to 90.7%, and in a U-shaped maze from 44% to 94%, with both reaching 100% when replanning mid-route, while a simple planner using the straightened map approached heavier planners and ran about 10 times faster. 1
- Hamel Husain said Quail, an open-source AI-SQL engine built with Modal, plans query planning together with LLM inference and can reach over 1 billion input tokens per minute per query on a single H100; he called the performance data impressive and the project still early. 1 He also argued that excellent human agents with AI skills, taste and initiative can change a business more than AI agents can, that finding such people offers "unreasonable alpha," and that this may be truer now than ever. 2 He said Jev and LLM judges are both classifiers and should be validated against trusted labels. 3
4. Product and Platform Updates
- Aravind Srinivas said a single prompt on Perplexity Computer (High Effort) can generate a Matterport-like effect, with a reposted item saying Computer can turn a property listing into an interactive, explorable 3D website. 1 He also said Computer can extract information from scientific sources such as Wiley Online Library to generate explainer videos, and that Wiley data is free for all Computer users with no setup or existing subscription required. 2
- Musk reposted SpaceXAI's Grok Bot Creator Rewards, which pays every two weeks based on number of users and frequency of use; templates must be posted on X with a link at least twice per payment cycle and X Money must be enabled, with payments initially following X original-content rewards. A poster said their Home Robots template earned $500 in two weeks. 1 Musk also reposted a Grok Bot that scans official sources, checks eligibility, finds proof, fills out a complete application and prepares it for submission to help users find claimable funds. 2
- NVIDIA AI said it worked with Meta to optimize the Muse Realtime Avatar model so avatars can keep pace in real-time conversation, turning Muse Realtime Voice into expressive, interactive avatars first used in Muse's real-time conversations. 1 NVIDIA also said agents may run for days, repeatedly call tools and hit retries, so safety policies must remain effective throughout, which is why it emphasizes runtime safety enforcement. 2 Jensen Huang, speaking on CNBC, said AI agents need clear behavioral boundaries and that those limits must remain effective throughout their work. 3
- Musk reposted a comparison of SpaceX capital efficiency: Blue Origin raised $30 billion while SpaceX used about $10 billion in private equity to complete orbital reusable booster development, roughly 650 flights and landings, over 10 million Starlink users, about 60 people flown to the ISS on Dragon, Starship full reusability about 95% complete, and the acquisition of xAI to build Starmind. 1 He also reposted that Falcon 9 and Dragon were at Florida's Launch Complex 40 preparing for NASA's Crew-13 mission to the ISS. 2
