infrastructure
Picking the right model per request beat always using the biggest one News
A new routing framework that chooses a different model for each request outperformed the strongest single fixed model by 14.6 percent, partly because the largest model gets many cheap questions wrong.
Google's private AI runs on sealed hardware, not on encrypted math News
Google's shipping private inference product runs Gemini inside hardware enclaves on custom chips, which is confidential computing rather than homomorphic encryption, and the company's actual homomorphic work is an unsupported research compiler.
Encrypted inference: can a model answer a question it cannot read? Lesson
The two competing ways to run AI on data the server is not supposed to see: sealed hardware enclaves, which ship today and are fast, and homomorphic encryption, which is mathematically stronger and still far too slow.
NVIDIA lines up six financiers to mobilize 500 billion dollars News
NVIDIA announced agreements with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to build independent financing platforms intended to mobilize more than 500 billion dollars of third-party capital for AI compute infrastructure.
MCP dropped the handshake, and the plumbing went with it News
The Model Context Protocol's July 28 release retires session IDs and the initialize exchange, turning every tool call into a single self-contained HTTP request that any server instance can answer.
Virginia orders Dominion to bill data centers for the power lines built to serve them News
The Virginia State Corporation Commission ordered Dominion Energy on July 31 to design, within 90 days, a tariff that charges large data-center customers directly for the substations and transmission lines built solely for them, rather than spreading those costs across every ratepayer.
Nashville voted 27-5 to take a data-center site by eminent domain News
Nashville's Metro Council gave final approval on August 5 to acquiring a 23-acre South Nashville property next to the city zoo by eminent domain if negotiation fails, blocking a planned data center that developer DC Blox bought for 23 million dollars in July.
Approximate nearest neighbor search: how a vector database finds a needle in a billion haystacks Lesson
Approximate nearest neighbor search is the algorithm that makes vector search fast enough to be useful - it finds the closest matches to a query embedding without comparing it against every item in the database, trading a small, tunable amount of accuracy for speedups of a hundred times or more. Every vector database and every retrieval-augmented system runs on it, and the accuracy it gives up is the hidden knob behind a lot of 'the retrieval just missed it' bugs.
OpenAI Rebuilt Voice So the Model Itself Decides When to Talk News
OpenAI's engineering posts on GPT-Live describe removing the separate turn detector from the audio path entirely and cutting session startup from six network round trips to one, treating a voice conversation as a live media system rather than a model feature.
EPA Says an Off-Grid Plant Built for One Data Center Escapes the Acid Rain Program News
An EPA guidance memorandum states that a fossil-fuel power plant with no physical connection to the utility grid, built to serve only an adjacent private data center, falls outside the federal Acid Rain Program and its permit, allowance, and monitoring requirements.
How a model is stored: safetensors, GGUF, and why one model arrives in 96 files Lesson
A trained model is just a large dictionary of numbered arrays saved to disk, and the file format that holds them determines whether the model loads safely, loads fast, and loads at all on your hardware.
NVIDIA Is Reportedly in Talks to Guarantee $250 Billion of OpenAI's Ohio Buildout News
The Wall Street Journal reports NVIDIA is discussing a roughly $250 billion credit guarantee for the lease and construction debt behind OpenAI's planned 10-gigawatt Ohio campus, a backstop that reportedly excludes the chips themselves.
Cloudflare now lets any site allow search crawlers while blocking AI agents and training bots separately News
Cloudflare has made three independently configurable AI crawler categories - Search, Agent and Training - available to every customer, and from September 15 new domains will block Agent and Training traffic on ad-bearing pages by default.
Distributed training: how one model gets split across thousands of chips Lesson
No single chip can hold a frontier model, so training is split across thousands of them in four distinct ways -- by data, by layer, by tensor, and by expert -- and choosing the right mix is what separates a cluster running at a third of its potential from one running at a tenth.
AI Revenue Now Covers the Data-Center Depreciation Bill, But Not the Full Cost News
A modeled industry report finds AI revenue first exceeded AI-infrastructure depreciation in late 2025, but the coverage is thin, excludes operating costs, and depends heavily on how long the chips last.
The '$1.65tn hidden AI debt' story, checked against the actual filings News
A Nikkei Asia estimate that five US tech giants carry $1.65 trillion in off-balance-sheet AI-related obligations is real as an estimate and grounded in verifiable filings of forward leases and purchase commitments, but it is not a hidden or auditable debt total, and much of the buildout's risk has been shifted to private-credit investors through project-finance vehicles.
Masayoshi Son Says the AI Economy Will Need $5 Trillion a Year by 2040 News
SoftBank's Masayoshi Son told his SoftBank World keynote that by 2040 AI will absorb about 20 percent of global GDP and require roughly $5 trillion a year in infrastructure investment, calling anyone who thinks AI is a bubble foolish.
New York just froze new hyperscale data centers for a year News
Governor Kathy Hochul signed an executive order creating what her office calls the nation's first statewide moratorium on new hyperscale data centers, halting discretionary state environmental permits for up to a year while New York writes new development standards.
Irish data centers now eat 23% of the country's electricity -- more than every city home combined News
Ireland's Central Statistics Office reported that data centers consumed 23% of the country's metered electricity in 2025, up from 5% in 2015, using more power than all urban households combined and rising even during a moratorium on new grid connections.
Mesh LLM lets you run models too big for any single machine by splitting them across peers News
Mesh LLM, the top project on Hacker News this week, runs models larger than any one machine can hold by partitioning them across networked peers -- layers 0-15 on one node, 16-31 on another -- over a serverless peer-to-peer transport, exposing a standard OpenAI-compatible API on localhost.
NVIDIA Starts Taking a Share of Its Cloud Partners' Revenue, Not Just Selling Them Chips News
NVIDIA announced a new arrangement where it earns a share of the cloud revenue its partners generate from NVIDIA-supported data center capacity, on top of its usual chip sales, with first partners Sharon AI and Firmus building campuses totaling hundreds of thousands of GPUs.
South Korea bets over a trillion dollars on chips, data centers, and robots News
The government and its biggest companies committed more than $1 trillion to memory fabs, AI data centers, and a goal of building tens of thousands of humanoid robots a year by 2028.
Big Tech is set to spend up to three-quarters of a trillion dollars on AI in 2026 News
Projected AI infrastructure spending for 2026 runs into the hundreds of billions, financed increasingly with debt, as OpenAI also moves into custom chips to cut inference costs.
What should an AI agent remember about you, and what leaks when it does? News
Researchers are asking whether AI agents are ready for real long-term memory, just as another study shows how much an agent's memory can quietly give away about the people it served.
Training vs inference: the two very different jobs inside every AI Lesson
Why building an AI model and using it are separate worlds with separate costs, and why that split explains custom chips, model prices, and where the real money in AI actually goes.
The quiet race to turn messy documents into AI-ready text News
Mistral released a new document-reading model the same week an open-source rival surged, both chasing the unglamorous job that quietly decides how well AI can read your files.
Qualcomm buys the software that lets AI run anywhere News
Qualcomm is paying about $3.9 billion for Modular, the Mojo language, and legendary compiler engineer Chris Lattner.
OpenAI designs its own chip to run its models News
With Broadcom, OpenAI unveiled a custom chip built for one job: serving its AI models cheaply.
NVIDIA's warm-water fix for AI's thirsty data centers News
A new NVIDIA cooling design claims to use almost no water inside the data center, though critics say that's only part of AI's water bill.
vLLM v0.23.0 Tool
The widely-used open engine for serving language models fast and cheaply. The latest release adds smarter memory handling for long conversations and faster GPU execution.
vLLM Tool
The popular open engine for serving AI models fast and efficiently when you need to handle real traffic.
slime Tool
The open-source large-scale asynchronous training framework from THUDM that Z.ai used to run the post-training scaling behind GLM-5.3.
exe.dev Tool
Persistent Linux virtual machines built for AI agents, with root access, SSH, a public hostname, a real network stack, and secrets injected by a host-side proxy rather than handed to the agent. Priced two ways: pooled capacity for steady workloads and per-second usage billing for bursty ones.
Vercel AI Gateway (Ling-3.0-flash) Tool
Vercel added Ling-3.0-flash to its AI Gateway with bring-your-own-key support and failover routing, free through August 3. Useful if you want the model behind a single gateway alongside other providers rather than wiring a second API.
SGLang v0.5.13 Tool
A high-performance open serving engine for language models. The new version turns on faster 'guess-ahead' decoding by default and trims scheduling overhead for quicker responses.
NeMo Switchyard Tool
NVIDIA's library for routing each task in a multi-model system to the model best suited to it, so a frontier model handles planning while a cheaper one handles execution. Shipped alongside Nemotron 3.5 Lightning as the connective tissue for mixed-model agent stacks.
Modular MAX + Mojo Tool
A programming language (Mojo) and compiler/runtime (MAX) for running AI models efficiently across different hardware instead of being locked to one chip vendor; now being acquired by Qualcomm but still openly available to developers.
Mistral OCR 4 Tool
A hosted document-reading model that converts scanned pages, PDFs, and complex layouts into clean structured text ready for a language model. Send a document, get back tidy text with the structure preserved.
MinerU Tool
Open-source tool that converts complex PDFs and office files into clean markdown and structured data that AI models can read reliably. Run it yourself for free, with nothing leaving your machine.
FastPLAID Tool
A Rust engine for multi-vector search, the indexing layer that makes late-interaction retrieval fast enough to serve in production instead of only benchmarking well.
AWS Agent Toolkit for AWS Tool
Official AWS-supported set of MCP servers, skills, and plugins for building AI agents that work with Amazon's cloud services, maintained by AWS itself.