MONOLITH COMPUTE · SANTA CLARA, CA
Epoch One puts 22,784 lattice cores, 36 GB of ECC GDDR7 and 1.92 TB/s behind a single driver — the same driver for the render farm and the LAN party. 450 watts. Two slots. No apologies.
SHIPS DECEMBER 2025 · 5-YEAR WARRANTY · 60 ISV CERTIFICATIONS AND COUNTING
EVERY PART SERVICEABLE WITH ONE #2 PHILLIPS · NO PULL-UPS, NO GLUE
Every GPU on the shelf today has two personalities. We built one that doesn't.
Pick up any brochure and you'll meet the split: a "creator" driver and a "gamer" driver, a consumer card and a workstation card — often the same silicon wearing two price tags, three drivers apart. The industry calls that segmentation. We call it homework for the customer.
Epoch One is our refusal. One driver, compiled once, installed in nine seconds, updated without rebooting — it doesn't know or care whether your kernel is a game, a viewport, or a diffusion model. It just runs it. At 128.5 TFLOPS. On the same card, in the same slot, for the engineer and the enthusiast alike.
We considered chiplets. Then we remembered what latency does to a viewport. BASALT-1 is one piece of TSMC N3P — 178 streaming multiprocessors, each carrying 128 ALUs and four third-generation matrix engines, wired to 128 MB of coherent last-level cache that answers in 8 TB/s.
Around the die: twelve 3 GB GDDR7 modules at 40 Gbps — the first shipping at that rate — on a 384-bit bus with link ECC. 1.92 TB/s, sustained, without a single re-fetch of your working set. Your dataset stops being a logistics problem.
Quiet is a spec. The vapor chamber pulls 340 W off the die and the GDDR7; six 8 mm composite heatpipes carry it into 412 g of skived fins; two 110 mm fluid-dynamic fans move 71 CFM at 1,850 rpm — and stop entirely below 60 °C. At full board power you will hear your thoughts.
Everything is serviceable. One Phillips driver, ten minutes, zero adhesive. The teardown slider above is not a metaphor — the card actually comes apart like that.
FIRST SILICON — MARCH 2025.
TSMC FAB 18, TAINAN · 300 MM · 52 GOOD DIES PER WAFER · 61% YIELD AT 748 MM²
Real cards, published figures, our lab. Where we win, it's marked. Where we don't — that's marked too.
RTX 6000 ADA WINS THIS ONE — A 300 W CARD TUNED FOR EFFICIENCY IS HARD TO BEAT. WE EXPLAIN IN 05 →
Kal is our kernel language: eleven statements, no memory model to memorize. The compiler tiles, schedules and binds to all 178 SMs — you keep the algorithm. And if you have a decade of CUDA, keep that too.
There is no Studio driver and no Game Ready driver. There is the driver. It hot-updates mid-render; your contexts survive, your clocks don't blink, and the rollback is one command if you don't like what you see.
mck importPoint it at a CUDA tree. 87% of the kernels we tested — including most of PyTorch's — transpile unmodified. The rest come back as a diff written in plain English, with the risky lines annotated by the compiler, not a wiki.
The driver journals every submission. Record a run once, then replay it instruction-exact, forever. The debugger steps backwards through time — because the hardware kept a timeline of it.
Binaries built for Epoch One run on Epoch Two, unmodified. We version the ISA so you don't have to re-schedule your quarter. This is a promise in the EULA, not a slide.
NumPy arrays, PyTorch tensors and Kal tensors share one heap on the card. .to('host') is the only copy you will ever write — and at 1.92 TB/s, 64 MB takes 34 µs.
A card that wins every row of every table is a card somebody made up. Here's where Epoch One loses — and why we'd lose the same way twice.
36 GB against the workstation cards' 48. We spent the power and the bill of materials on bandwidth and price instead of capacity. If capacity is your constraint, the RTX 6000 Ada is the right tool and its 48 GB is glorious. For everyone else: Epoch One Max, 96 GB, ships H1 2026 on the same driver and the same kernel binaries.
The RTX 5090 posts a higher sparse-FP8 number and we congratulate it. We optimized for dense FP16 — 578.6 TFLOPS — where production inference actually lives. If your datacenter is billed in sparse-FP8 sprints, buy Blackwell.
FP64 runs at 1/128 rate — about 1.0 TFLOPS against Ada's 1.42 and the 5090's 1.6. We chose not to spend die area there. If your day is spent in FP64, this is genuinely not your card, and we're at peace with that.
If the card's only job is frames, the RTX 5090 is $500 less and it is magnificent. Buy it. We said so in the table, in print, with a diamond next to its price.
60 ISV certifications against NVIDIA's 200+. Every major DCC is covered; the long tail isn't. The list grows monthly and the queue is public — we don't do "certified-ish."
IF YOU NEED ONE OF THESE THINGS TO BE DIFFERENT, BUY THE OTHER CARD. THEY'RE GOOD CARDS. WE PUT THAT IN WRITING TOO.
FULLY REFUNDABLE UNTIL YOUR UNIT SHIPS · LIMIT TWO PER HOUSEHOLD — WE'RE NOT DOING SCALPER MATH