Ground Truth.
AI, checked against the source.

← All topics

hardware

Everything on Ground Truth tagged “hardware” — 38 items.

A 2.8-trillion-parameter model runs on a laptop by streaming weights off four SSDs News

A project called Deltafin runs the full uncompressed Kimi K3 model — 2.8 trillion parameters, with 1.45 TB of expert weights — on a single MacBook Pro by streaming experts from four external SSDs on demand, sustaining almost exactly one token per second.

Cerebras is serving an open 27B model at 1,500 tokens a second, and the free tier caps it exactly News

Cerebras now serves Qwen 3.8 27B at roughly 1,500 output tokens per second, but its own rate-limit page caps free-tier users at 90,000 tokens per minute -- almost precisely the model's raw output rate -- so the headline speed only becomes usable on the paid tier.

DRAM contract prices nearly doubled in a single quarter News

Conventional memory contract prices rose roughly 93% to 98% quarter over quarter in early 2026 and are forecast to climb another 58% to 63%, as suppliers divert capacity to AI servers -- repricing the exact component local AI depends on.

OpenAI publishes first Jalapeno results, claiming up to 1.9x more work per watt than the systems it tested against News

OpenAI released measured results for Jalapeno, its Broadcom-co-designed inference chip, reporting 1.5-1.9x more AI work per watt, 1.7-3.6x lower latency, and 2.1-4.1x higher performance on interactive workloads, with kernels its own model wrote.

Apple's Mac Studio now holds 512GB of unified memory, which solves capacity and leaves speed exactly where it was News

The M5 Ultra Mac Studio configures to 512GB of unified memory at 1.2TB/s, enough to load almost any open-weight model in existence, but its memory bandwidth still sets a hard ceiling on how fast those models can generate text.

LLMs are less resilient to bit flips than accuracy suggests News

A supercomputing-conference study that injected more than 13 million simulated hardware faults into language model inference found that benchmark accuracy barely moves while the quality of generated text degrades badly, and that 4-bit quantized models are more robust than full-precision ones.

A llama.cpp fork is reviving $200 AMD cards nobody else supports News

A specialist fork of llama.cpp ships hand-written kernels for AMD's decade-old GFX906 architecture, making cheap used MI50 and Radeon VII cards usable for local inference, and upstream maintainers are now discussing porting the work back.

Etched raised 700 million dollars and shipped its first rack to Jane Street News

Etched announced on August 18 that it shipped its first inference rack to Jane Street and raised 700 million dollars at a 21 billion dollar valuation, betting that frontier inference belongs on hardware co-designed for it rather than on general-purpose accelerators.

Unitree lists in Shanghai as the rare profitable humanoid robot maker News

Unitree Robotics shares begin trading on Shanghai's STAR Market on August 19 under ticker 688836, and its filing shows 1.70 billion yuan of 2025 revenue and 278 million yuan of net profit against comparable listed robot companies that the prospectus says are all loss-making.

OpenAI put its most intelligent model on Cerebras chips at 750 tokens a second News

OpenAI is previewing Ultrafast, a service tier that runs GPT-5.6 Sol on Cerebras hardware at up to 14 times the speed of standard processing and up to 750 output tokens per second.

China's biggest memory maker is booked through 2027 News

ChangXin Memory Technologies has reportedly sold out its DRAM output through the end of 2027 as PC brands rushed to secure supply, and consumer memory prices have stayed near their highs since.

The Full 2.8-Trillion-Parameter Kimi K3 Now Runs on Sixteen Desktop Boxes News

An operator has the complete Kimi K3 checkpoint running across sixteen GB10 mini-workstations wired through a single 400G switch, producing roughly 21 to 25 tokens per second for one user, on hardware with a verifiable floor around $57,200.

High Bandwidth Flash Became a Spec Today, Not a Product You Can Buy News

SK hynix and Sandisk published the first standard for High Bandwidth Flash at FMS 2026, defining a NAND memory tier of up to 512GB per stack with a top bandwidth grade near three terabytes a second - with no named accelerator, price or availability date.

The Cheap 284B Rig Is Really 768GB of Server Memory News

A builder running DeepSeek V4-Flash at 33 tokens a second on two RTX 3090s is holding about 6.6GB of weights per card and roughly 170GB per instance in system memory on a four-socket enterprise server, which is where the model actually lives.

The "2x GB200 bandwidth" Chinese chip claim is a 2027 projection, and the arithmetic gives 1.67x News

A widely shared claim that a Chinese accelerator delivers twice the memory bandwidth of NVIDIA's GB200 traces to a roadmap part expected in early 2027, compared 64-at-a-time against a full NVIDIA rack, and the published numbers work out to 1.67 times at rack level while the single chip lands below a shipping GB200.

The FCC just added every foreign-made advanced robot to its national security Covered List News

On July 28 the FCC added all foreign-produced advanced robotic devices and foreign-produced power inverters to its Covered List, blocking them from new equipment authorizations, on national security determinations that cite remote commandeering and surveillance risk rather than naming any country or company.

Kimi K3 is downloadable, but the floor to run it is eight datacenter GPUs News

Kimi K3's 1.56-terabyte checkpoint needs a single eight-GPU B300 or MI355X node as its practical minimum, and no version of llama.cpp can load it today, so open weights currently mean operator-scale rather than local.

Why AI Inference Runs Out of Memory Bandwidth Before It Runs Out of Math Lesson

Generating text with a language model is limited by how fast weights can be moved from memory into the processor, not by how fast the processor can multiply, which is why most of a GPU sits idle during inference.

Mixed-precision training: why models are trained in half-broken numbers on purpose Lesson

Modern models are trained using 16-bit and even 8-bit numbers instead of the 32-bit standard, roughly doubling speed and halving memory, by carefully keeping full precision exactly where the arithmetic would otherwise fall apart.

Anthropic plans up to two gigawatts of AMD chips, with AMD committing up to $5 billion back News

AMD said Anthropic plans to deploy up to 2 gigawatts of MI450-series capacity starting in the first half of 2027, and that AMD has committed to a future equity investment of up to $5 billion in Anthropic.

AMD and Cerebras split AI inference across two different chips News

AMD and Cerebras announced a joint inference offering on July 23 in which AMD's Helios racks process the prompt and Cerebras's wafer-scale engine generates the tokens, claiming up to five times the tokens per watt of a Cerebras-only setup.

A Huawei-chip training report shows what leaving CUDA actually costs News

SLAI's technical report documents full-parameter post-training of a DeepSeek-V4 model on Huawei Ascend hardware, and the work list -- rebuilt collectives, converted checkpoints, hand-written kernels -- is the real measure of chip independence.

Google's two opposite bets: a Gemini-specialized chip and an EU order to open Android AI News

Google is reportedly designing a server chip called Frozen v2 that hardwires Gemini's architecture for six-to-ten times more tokens per watt, even as the European Commission adopted binding measures forcing Android to open eleven AI capabilities to rival assistants, making Google simultaneously bet on locking Gemini into silicon and being forced to unlock Gemini's Android advantages.

Google Falls Off One Leaderboard's Top 15, as a Report Describes a Gemini-Specific Chip News

Google dropped out of the top 15 on LLM Stats' composite leaderboard while remaining its fastest model, and Reuters separately reported an unannounced Gemini-specific inference chip.

OpenAI is selling a $230 keyboard with a dial for how hard the AI thinks News

OpenAI has launched Codex Micro, a $230 mechanical control deck built with accessory maker Work Louder that puts agent status on RGB keys and reasoning effort on a physical rotary dial.

DeepSeek is designing its own AI chip -- and raising outside money for the first time News

Chinese AI startup DeepSeek is developing its own chip aimed at running trained models rather than training them, and is simultaneously raising its first-ever outside capital -- about $7 billion at a $52-59 billion valuation.

SK Hynix's Nasdaq listing raises $26.5 billion, the largest first-time US listing by a foreign company ever News

SK Hynix raised $26.5 billion in a Nasdaq listing that topped Alibaba's record for the largest US IPO ever by a foreign company, priced on the strength of its dominance in the high-bandwidth memory that AI accelerators depend on.

Meituan open-sources LongCat-2.0, a trillion-parameter model it says was trained end-to-end on Chinese chips News

Meituan released LongCat-2.0, a 1.6-trillion-parameter open-weight (MIT) model that ran anonymously as 'Owl Alpha' for two months and was, the company says, both trained and served entirely on domestic Chinese AI ASICs with no Nvidia GPUs.

Apple sues OpenAI, alleging it poached staff and stole hardware secrets to build AI devices News

Apple filed suit against OpenAI in federal court on July 10, 2026, alleging former Apple employees now at OpenAI directed current staff to hand over unreleased-device secrets and that one ex-employee downloaded confidential files after leaving.

Anthropic and UST put Claude Code to work validating computer chips News

Anthropic and IT services firm UST announced a 'Physical AI' alliance using Claude Code to read chip schematics and pinouts and auto-generate regression tests on UST's iDEC platform, which the companies say cuts hardware validation cycle times by 50 to 70 percent.

The AI Memory Boom Just Made Your Next Laptop Much More Expensive News

Apple raised prices across its lineup, with a top MacBook Pro reaching $10,000, because AI data centers are consuming so many memory chips that the price of RAM has quadrupled this year.

AI is learning a 'dark art' that even expert engineers struggle with News

Designing radio-frequency chips is so reliant on hard-won physical intuition that engineers call it a dark art - and now AI is starting to do it, a sign the automation frontier is moving into deep specialist craft.

Training vs inference: the two very different jobs inside every AI Lesson

Why building an AI model and using it are separate worlds with separate costs, and why that split explains custom chips, model prices, and where the real money in AI actually goes.

OpenAI designs its own chip to run its models News

With Broadcom, OpenAI unveiled a custom chip built for one job: serving its AI models cheaply.

A robot that runs its own experiments — and sometimes fails when it matters News

NVIDIA researchers gave AI coding agents full control of a physical robot lab — including automated reset and vision-based success checking. One agent inserted a graphics card into a motherboard. The headline success rate is real but requires a close read.

SC25 LLM reliability assessment Tool

Fault-injection harness that flips individual bits during language model inference through PyTorch hooks, then restores them, so you can measure how your own model degrades under simulated soft errors instead of assuming it is resilient.

Copperhead Tool

An open-source AI agent for KiCad electronics design that drew wide attention on Hacker News this week. Apache 2.0.

Codex Micro Tool

A $230 mechanical control deck for driving OpenAI's Codex agents, built with keyboard maker Work Louder. 13 switches, a joystick, a touch sensor, RGB keys showing live agent status, and a rotary dial that adjusts reasoning effort -- turning an API parameter into a physical knob. Nothing it does is impossible with keyboard shortcuts; the pitch is ambient awareness when supervising several agents at once.