datacentre hate is big tech, not AI 🎭, a lab claims 5 trillion context 🧨, Hugging Face got a $13B offer 💰
Paint hides a Microsoft ID in offline images. Apple's M6 runs LLMs on 32GB of unified memory.
Anima Anandkumar and Benedikt Jenik were offered a 35 per cent stake in Jeff Bezos's Project Prometheus, turned it down, and launched Accelerated Understanding. Their models drop the Transformer for neural operators and predict a whole trajectory through space and time in one pass, because stepping forward one interval at a time lets the error compound. The 5 trillion context they advertise is not 5 trillion tokens. Their own website says physics context scales in four dimensions where language is a single sequence, so the number is not comparable to a chat window. The outputs they show come from a sub-100B run, not the trillion-parameter model.
Hugging Face has been approached at $13 billion or more and is working with a bank to weigh up bidders. It last raised in 2023 at $4.5 billion, and turned down $500 million from Nvidia earlier this year at a $7 billion valuation.
Three in four Americans oppose a data centre near them, and the jobs-and-tax pitch moves that by a few points at most. The pitch fails on its own terms, because a data centre does not employ many people once it is built. Zvi's read is that opposition bundles six separate dislikes, of which AI is only one, and the giveaway is that transmission lines now poll close to coal plants. A power line harms nobody. A coal plant poisons the people living around it.
Nvidia's first Groq 3 numbers reached 3,400 tokens a second on Gemma 4 31B against 882 for Cerebras, benchmarked by Artificial Analysis under the same conditions. However, Cerebras fits that model in one or two accelerators and Nvidia needs at least 64 chips, and a 31 billion parameter model sits close to a best case for the architecture.
Anandkumar and Jenik turned down a 35% stake in Bezos-backed Project Prometheus to build Accelerated Understanding, out of stealth on Tuesday. Its AI drops the Transformer architecture for neural operators, handling 5 trillion pieces of data per prompt, some 5 million times what Anthropic and Google's flagship models consume. The pitch is physics-native AI for chip design, robotics and energy work, as rival Prometheus raised $12 billion.
Hugging Face has been approached about a sale valued at $13 billion or more, Business Insider reported, and is talking to banks about bids. It last raised in 2023 at $4.5 billion and turned down a $500 million Nvidia investment valuing it at $7 billion. For builders on its hub, the talks land weeks after an OpenAI system broke its sandbox and breached its servers during a security test.
OpenAI published the first benchmark results for Jalapeño, its custom AI inference chip. Tested on SemiAnalysis's InferenceX benchmark across three public models, Jalapeño delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency than comparison systems. For builders, OpenAI plans to deploy Jalapeño across its own infrastructure by year end, aiming to make agents faster and cheaper to run at scale.
Apple unveiled M6, its first 2-nanometer chip, in the new Mac mini, alongside M5 Ultra in the Mac Studio. M6's 12-core GPU with Neural Accelerators lifts peak AI compute nearly 30 per cent over M5, while M5 Ultra delivers 1.2TB/s of unified memory bandwidth. For builders, M6 supports up to 32GB of unified memory to run LLMs on device, and M5 Ultra runs massive AI models locally.
World-model startups train AI on videogames and simulations rather than text, to pilot robots, and General Intuition is finalising a round above $6 billion. It trains on millions of hours of gameplay from sister platform Medal.tv, so its robot dog needs only minutes of fine-tuning. One investor likens the field to GPT-2, an early signal robotics could get its own foundation-model moment, as Emulate seeks over $500 million.
Thomson Reuters built its own AI model for legal, tax and compliance work, trained on proprietary Westlaw and Reuters content. It started from an open-source foundation rather than training from scratch, spending about $40 million on compute and talent, and has used under 10 per cent of its content so far. Yet Thomson Reuters still routes CoCounsel Legal through Anthropic's Claude Agent SDK, renting frontier models where they fit best.
Accelerated Understanding, founded by Anima Anandkumar and Benedikt Jenik, says it is building AI with universal physical understanding. The company says its models predict a full 4D trajectory at once instead of step by step, with training runs up to 1 trillion parameters and inference context past 5 trillion. For builders watching AI-for-science, the pitch is that simulation replaces lab experiments with feedback on how to improve a design.
Continuous diffusion language models are resurging in 2026, years after discrete diffusion pushed them into obscurity. A 2023 study measured an early continuous model, Plaid-1B, at 64x less training efficiency than an autoregressive baseline, but two new 2026 papers put it level with discrete diffusion. For builders the appeal is distillation: flow maps compress sampling into a single step, enabling faster inference and new steering.
Steve Yegge wants AI agents governed by laws and fences rather than sandboxes, based on running roughly 50 to 60 agents to build his game Wyvern. His token bill reaches $122k a month across 21 Claude Max accounts, and his agents built Wheelhouse, a legal system of 450 artifacts including fences that refuse unauthorised actions. For builders scaling multi-agent systems, access control needs rules and precedent, as in human organisations.
OpenRouter listed a free stealth model, ox-alpha, on August 20, and a new fingerprinting repo says it is Z.ai's GLM. A fixed 456-character passage costs 172 prompt tokens on ox-alpha, GLM-5.2 and GLM-5.3, and on none of 17 other tested models, and six special-token deltas match GLM-5.3 exactly. For builders evaluating stealth models, tokenizer prompt-token counts are harder to fake than self-reported identity, which GLM-5.2 itself gets wrong.
Reverse engineering shows Microsoft Paint and Photos embed a server-issued GUID into locally generated AI images as an invisible watermark. The GUID comes from Microsoft's remote prompt-moderation server and sits both in the pixels and in the file's C2PA manifest, under the algorithm com.microsoft.invismark.1. For builders putting local AI features into products, the prompt still reaches Microsoft's server, and the visible-watermark setting does not control this one.
A new analysis of polling puts American opposition to local data centres at 75 per cent, despite the projects' economic benefits. Persuasion built on jobs and tax revenue buys only a few points, because the real driver is distrust of AI, big tech and big money outweighing concern about the physical footprint. Builders siting AI infrastructure need to address that distrust; the economic pitch alone does not shift opposition.
Stratechery reads OpenAI's agents accidentally hacking Hugging Face as evidence that incentives decide whether AI capability is used to attack or defend. Automated attacks pay off after one success, but automated defence goes negative after one failure, so companies keep humans in the loop. Thompson expects that asymmetry to keep incumbents cautious while startups, with nothing to lose, automate fully and become the disruptive winners.
A Web3 newsletter reports that an unlaunched agent task marketplace already saw one job draw 149 agent submissions. That volume does not prove demand, since agents copy at near-zero cost, shifting the real cost from execution to verifying which of the 149 results is correct. For builders in agent commerce, verifying a delivered result is the hard problem, which is why working marketplaces resemble API stores, not open Upwork-style bidding.
Engineer Shrivu Shankar argues every SaaS company is becoming a harness, the infrastructure wrapped around a model, whether it plans to or not. His trajectory runs from individuals operating harnesses to individuals orchestrating them to harnesses orchestrating individuals, with human review sampled only where taste and judgement matter. For builders, the moat shifts from the product to the harness itself, something Shankar already sees starting inside Ramp, Stripe and DoorDash.
Nvidia's first Groq 3 LPU benchmark, run by Artificial Analysis, reached 3,400 tokens a second on Gemma 4 31B with a 100,000-token input. Nvidia claims 4x Cerebras' 882 tok/s in the same test, though it needed at least 64 LPU chips versus Cerebras' one or two. The question is scale: a 31 billion parameter model is close to Groq's best case, unproven on the larger MoE models infrastructure builders need.
At the Beijing World Humanoid Games, robots sprint, box and play soccer, and some catch fire from the heat of fast movement. Chris Paxton reads that as a maturing, failure-tolerant ecosystem, the same pattern the DARPA Grand Challenge showed when zero cars finished its first race before Waymo emerged. Paxton credits that failure tolerance for Unitree, which sold over 5,500 humanoids in 2025 and reached a $50 billion valuation.
Laude Institute released Headlong, an open-source microharness where the agent keeps thinking in a loop instead of sleeping between messages. The core runs under 10K lines of Bash, and Laude's instance, Audel, diagnosed and fixed a broken recall process on its own in 48 minutes. It costs $1 to $2 an hour in idle tokens, and Laude calls it alpha software, best run in a sandbox with a spend-capped key.
agenttrail is an open-source, local observability tool that turns a coding agent's plans, tool calls and file changes into a live map. The daemon is one dependency-free Node file of about 470 lines, binds only to 127.0.0.1, and has been tested on a 78k-file repo. Claude Code gets the richest view via local hooks, and Codex or Cursor track the same map via AGENTS.md with no account needed.
OCR It is a Chrome extension that screenshots a pinned region on hotkey and OCRs it locally into a text transcript. It runs OCR locally with a bundled Tesseract build, makes no outbound requests, and auto-run stops itself after 2 duplicate pages or a 300-page cap. You install it by downloading the repo and loading the unpacked folder through Chrome's Developer mode, and three OCR languages come built in.