Key Takeaways
- Hugging Face open-sourced a 1T parameter model achieving 1000+ tokens/s with FP4 quantization and DFlash, and MiniMax released M3, a 428B-parameter multilingual model with 23B active parameters.
- Perplexity integrated Deep Research as a native skill in Computer, based on a Search as Code architecture, and launched Plan Mode for all users.
- xAI launched the Grok Build Plugin Marketplace beta with MongoDB, Vercel, and Sentry plugins.
- NVIDIA Research open-sourced MotionBricks, a dataset of 350k+ motion clips enabling real-time character animation at 15,000 fps.
- Frontier models failed the Beninatto-Trombetti translation test, revealing they reorganize knowledge rather than truly understand.
- Small AI-driven quant funds are challenging efficient market hypotheses, potentially leading to concentrated trades and model poisoning risks.
1. Major Open-Source Model Releases
- Hugging Face open-sourced a 1T parameter model that achieves over 1000 tokens per second on a single node with 8 GPUs. The model employs FP4 quantization and DFlash technology for high-efficiency inference, and weights are openly available. — via 1
- MiniMax released the M3 open-source multilingual model, which has approximately 428B total parameters with about 23B active. It supports text, image, and video reasoning and is available on HuggingFace with free GPU-accelerated endpoints. — via 1 2
- NVIDIA Research open-sourced MotionBricks, containing over 350,000 motion clips that enable real-time character animation at 15,000 frames per second. The model requires no manual transitions or fine-tuning and is also applicable to robotics. — via 1 2
2. AI Agent and Tool Ecosystem
- Perplexity integrated Deep Research as a native skill into its Computer product, available for Pro and Max users. Deep Research is powered by a Search as Code architecture where the model writes code to execute thousands of parallel retrieval steps, outperforming traditional deep research methods in benchmarks. Additionally, Plan Mode was released to all Computer users, clarifying requirements via questions and connecting to the correct data sources before execution. — via 1 2 3
- xAI launched the Grok Build Plugin Marketplace in beta, featuring plugins for MongoDB, Vercel, and Sentry. Users can use the MongoDB plugin to explore data and build vector searches, the Vercel plugin to deploy and build apps, and the Sentry plugin to find and fix errors, all from the terminal. — via 1
- OpenAI updated Codex to allow saving rate limit resets for later use, providing one free reset to Go, Plus, Pro, and Business users. — via 1 2
3. Frontier Model Capabilities and Insights
- Frontier models failed the Beninatto-Trombetti translation test, which requires correctly updating a meta-linguistic statement within a sentence (e.g., translating "3 parole" to "4 words" after changing the text length). This demonstrates that models do not truly understand content but rather reorganize existing knowledge. — via 1
- Deedy reported that small groups running AI-driven quant funds are achieving high returns, with many doubling capital in months. This challenges efficient market hypotheses, leads to concentrated trades, and makes markets vulnerable to model-poisoning attacks, though the alpha is expected to decay eventually. — via 1
- A new predictive data debugging method can reveal and shape what a model will learn before training, avoiding costly irreversible runs. In DPO datasets, it found broken guardrails, hallucinations, and absurd content, underscoring the direct link between data quality and model quality. — via 1
