MONOLITH COMPUTE · SANTA CLARA, CA

Most GPUs are peripherals.
This one is an instrument.

Epoch One puts 22,784 lattice cores, 36 GB of ECC GDDR7 and 1.92 TB/s behind a single driver — the same driver for the render farm and the LAN party. 450 watts. Two slots. No apologies.

SHIPS DECEMBER 2025 · 5-YEAR WARRANTY · 60 ISV CERTIFICATIONS AND COUNTING

DRAG TO ROTATE · HOVER TO IDENTIFY PARTS
IDLE
BOARD12 W
CORE210 MHz
EDGE38 °C
FANS0 rpm

EVERY PART SERVICEABLE WITH ONE #2 PHILLIPS · NO PULL-UPS, NO GLUE

BASALT-1 · TSMC N3P 36 GB GDDR7 ECC · 40 GB/S-PIN 1,920 GB/S SUSTAINED PCIE 5.0 ×16 2-SLOT · 267 MM · 450 W 4× DP 2.1B + HDMI 2.1B
01THESIS

Every GPU on the shelf today has two personalities. We built one that doesn't.

Pick up any brochure and you'll meet the split: a "creator" driver and a "gamer" driver, a consumer card and a workstation card — often the same silicon wearing two price tags, three drivers apart. The industry calls that segmentation. We call it homework for the customer.

Epoch One is our refusal. One driver, compiled once, installed in nine seconds, updated without rebooting — it doesn't know or care whether your kernel is a game, a viewport, or a diffusion model. It just runs it. At 128.5 TFLOPS. On the same card, in the same slot, for the engineer and the enthusiast alike.

1 4 9
THE MARK'S PROPORTIONS ARE 1 : 4 : 9 — THE SQUARES OF ONE, TWO, THREE. THE LAST ARTIFACT HUMANITY FOUND IN THOSE PROPORTIONS WAS BURIED THREE MILLION YEARS BEFORE FIRE. WE PUT IT ON A GRAPHICS CARD. AMBITION IS A SPEC TOO.
0MM² DIE — BASALT-1
0BILLION TRANSISTORS
0MB STRATA CACHE
0GHZ BOOST CLOCK
0TB/S MEMORY BANDWIDTH
02SILICON

BASALT-1. 748 mm² of intent.

BASALT-1 · ENGINEERING SAMPLE ES2 · 748 MM²HOVER THE DIE OR THE LEGEND

We considered chiplets. Then we remembered what latency does to a viewport. BASALT-1 is one piece of TSMC N3P — 178 streaming multiprocessors, each carrying 128 ALUs and four third-generation matrix engines, wired to 128 MB of coherent last-level cache that answers in 8 TB/s.

Around the die: twelve 3 GB GDDR7 modules at 40 Gbps — the first shipping at that rate — on a 384-bit bus with link ECC. 1.92 TB/s, sustained, without a single re-fetch of your working set. Your dataset stops being a logistics problem.

    AIR IN · 71 CFM EXHAUST 110 MM FJD FANS ×2 SKIVED FINS · 412 G VAPOR CHAMBER · 340 W 14-LAYER PCB COMPOSITE HEATPIPE ×6 BASALT-1 · 12× GDDR7 BACKPLATE · 1.8 MM MONOLITH COMPUTE EPOCH ONE · SECTION A–A VERT SCALE ×2 · REV C SHEET 3 / 7 · 450 W QUIET IS A SPEC — 34 dBA AT FULL 450 W · FANS OFF BELOW 60 °C

    Quiet is a spec. The vapor chamber pulls 340 W off the die and the GDDR7; six 8 mm composite heatpipes carry it into 412 g of skived fins; two 110 mm fluid-dynamic fans move 71 CFM at 1,850 rpm — and stop entirely below 60 °C. At full board power you will hear your thoughts.

    Everything is serviceable. One Phillips driver, ten minutes, zero adhesive. The teardown slider above is not a metaphor — the card actually comes apart like that.

    FIRST SILICON — MARCH 2025.
    TSMC FAB 18, TAINAN · 300 MM · 52 GOOD DIES PER WAFER · 61% YIELD AT 748 MM²

    03NUMBERS

    First or second in every row but one.

    Real cards, published figures, our lab. Where we win, it's marked. Where we don't — that's marked too.

    PEAK FP32 TFLOPS · HIGHER IS BETTER

    Epoch One
    128.5
    RTX 5090
    104.8
    RTX 6000 Ada
    91.1
    Pro W7900
    61.0

    FP32 PER WATT TFLOPS / W · HIGHER IS BETTER

    Epoch One
    0.286
    RTX 6000 Ada
    0.304
    Pro W7900
    0.207
    RTX 5090
    0.182

    RTX 6000 ADA WINS THIS ONE — A 300 W CARD TUNED FOR EFFICIENCY IS HARD TO BEAT. WE EXPLAIN IN 05 →

    04PROGRAM

    You describe the math.
    We'll find the silicon.

    Kal is our kernel language: eleven statements, no memory model to memorize. The compiler tiles, schedules and binds to all 178 SMs — you keep the algorithm. And if you have a decade of CUDA, keep that too.

    
              
    
              
    
            
    MCC · BUILD LOG
    9 : 214LINES TO EXPRESS A DENSE GEMM — KAL VS CUDA C++
    87%OF TESTED CUDA KERNELS TRANSPILE UNMODIFIED VIA mck import
    1DRIVER. THERE IS NO SECOND DRIVER.

    One driver, forever

    There is no Studio driver and no Game Ready driver. There is the driver. It hot-updates mid-render; your contexts survive, your clocks don't blink, and the rollback is one command if you don't like what you see.

    mck import

    Point it at a CUDA tree. 87% of the kernels we tested — including most of PyTorch's — transpile unmodified. The rest come back as a diff written in plain English, with the risky lines annotated by the compiler, not a wiki.

    Deterministic replay

    The driver journals every submission. Record a run once, then replay it instruction-exact, forever. The debugger steps backwards through time — because the hardware kept a timeline of it.

    Forward-architecture guarantee

    Binaries built for Epoch One run on Epoch Two, unmodified. We version the ISA so you don't have to re-schedule your quarter. This is a promise in the EULA, not a slide.

    Zero-copy Python

    NumPy arrays, PyTorch tensors and Kal tensors share one heap on the card. .to('host') is the only copy you will ever write — and at 1.92 TB/s, 64 MB takes 34 µs.

    05HONESTY

    What it isn't.

    A card that wins every row of every table is a card somebody made up. Here's where Epoch One loses — and why we'd lose the same way twice.

    Not the biggest memory.

    36 GB against the workstation cards' 48. We spent the power and the bill of materials on bandwidth and price instead of capacity. If capacity is your constraint, the RTX 6000 Ada is the right tool and its 48 GB is glorious. For everyone else: Epoch One Max, 96 GB, ships H1 2026 on the same driver and the same kernel binaries.

    36 < 48 GB3RD OF 4, BY DESIGN

    Not the highest AI peak.

    The RTX 5090 posts a higher sparse-FP8 number and we congratulate it. We optimized for dense FP16 — 578.6 TFLOPS — where production inference actually lives. If your datacenter is billed in sparse-FP8 sprints, buy Blackwell.

    2,314 < 3,352TOPS · FP8 SPARSE

    Not a double-precision powerhouse.

    FP64 runs at 1/128 rate — about 1.0 TFLOPS against Ada's 1.42 and the 5090's 1.6. We chose not to spend die area there. If your day is spent in FP64, this is genuinely not your card, and we're at peace with that.

    1/128 RATE~1.0 TFLOPS FP64

    Not the cheapest way to game.

    If the card's only job is frames, the RTX 5090 is $500 less and it is magnificent. Buy it. We said so in the table, in print, with a diamond next to its price.

    $2,499 > $1,999MSRP, USD

    Not certified everywhere. Yet.

    60 ISV certifications against NVIDIA's 200+. Every major DCC is covered; the long tail isn't. The list grows monthly and the queue is public — we don't do "certified-ish."

    60 < 200+CERTIFIED APPLICATIONS

    IF YOU NEED ONE OF THESE THINGS TO BE DIFFERENT, BUY THE OTHER CARD. THEY'RE GOOD CARDS. WE PUT THAT IN WRITING TOO.

    06RESERVE

    Reserve.

    $2,499 USD · ONE CONFIG · NO "FOUNDING EDITION" TAX
    SHIPS DECEMBER 2025, WORLDWIDE, SAME DAY

    FULLY REFUNDABLE UNTIL YOUR UNIT SHIPS · LIMIT TWO PER HOUSEHOLD — WE'RE NOT DOING SCALPER MATH

    WHAT $2,499 BUYS

    • The whole cardNo cut-down memory, no locked features, no tiers
    • Fits your case2-slot · 267 mm · 2× 8-pin or 12V-2×6
    • One driverConsumer and professional, forever
    • 5-year warranty48-hour advance replacement, cross-shipped
    • Driver supportGuaranteed through 2035, in writing
    • Epoch Two trade-in40% credit — working or not