Nagare Labs
Nagare Labs · Atlanta

Every leap in computing has been a leap in integration. The next one is memory.

Transistor, integrated circuit, VLSI. Now 3D heterogeneous integration is bringing logic and memory together, and Nagare Labs is taking the next step: memory designed for AI infrastructure, integrated directly on the AI accelerator.

Fast enough to keep up with logic. Dense enough to hold the active model parameters and its KV cache. Built with the thermal and power envelope to operate on top of a multi-kilowatt accelerator.

Product · Where memory is going

The memory wall has one way out: up.

The most energy-efficient way to compute is to keep data next to compute and not move it. A single low-precision multiply-add costs about 10 fJ. Fetching the number it operates on from HBM costs 3–4 pJ, roughly 1,000× more, almost all of it spent hauling bits across the package. SRAM keeps data close, but runs out of capacity almost immediately.

Z2RAM is Nagare Labs' answer: a memory with SRAM-class speed at DRAM-class density, designed from the start to sit directly on the AI accelerator.

SRAM

Right next to compute, but tiny.

Bandwidth, latency, and energy per bit are all excellent. Capacity is not: six transistors per bit, half the die, and no longer shrinking.

HBM

Capacity at the end of a straw.

Intricately and expensively stacked beside the accelerator, connected by a few thousand lanes through an interposer. Every bit pays for the trip in energy, and every access pays for it in latency.

Stacked DRAM

Closer, but capped by heat.

Conventional DRAM bonded onto the accelerator recovers much of what SRAM offers. But DRAM has to stay below about 95 °C, which throttles the logic beneath it. So stacks stay short, most of the model goes back to the interposer, and each layer is a separate die to bond and yield.

Z2RAM

Fast, dense, and made for the top of the die.

A transistor-only cell built from oxide semiconductors, deposited in layers on one die. SRAM-class access, DRAM-class density, and stable at accelerator operating temperatures.

What memory on compute must doSRAMHBMStacked DRAMZ2RAMHow
Keep data next to computeNo interposer; hundreds of thousands of lanes, not thousandsMemory layers stacked directly over logic, with connections running vertically through every layer. Bits travel micrometers, not millimeters.
Well under a pJ per bitFree the power budget for computeNo bumps or bonds between memory layers at all, and almost no energy spent holding data in place.
Low latencySRAM-class access, no interface trainingA fast, non-destructive read at SRAM-class access times, and refresh so infrequent it stays out of the way of the workload.
Higher-temperature operationStable where conventional DRAM is notNear-zero-leakage oxide-semiconductor transistors hold data at temperatures where conventional DRAM cells cannot, and draw far less power doing it.
Scalable by designDeposited layers, not bonded diesZ2RAM adds capacity by depositing another memory layer on the same die and patterning them all together, the way 3D NAND grew. One thin die with low thermal resistance.
meetspartlydoes not
14×

Faster than DRAM

Access speed nearly on par with SRAM, so AI accelerators compute instead of wait.

projected at Gen 1
5×

Denser than SRAM

More memory in less space, for larger models and longer context, with a roadmap past 10× as layers are added.

projected at Gen 1
3×

Less read energy than SRAM

Lower power and heat, with no refresh cycles burning energy to hold data in place.

projected at Gen 1
Roadmap · density scales with layers

Bit density relative to 2 nm high-density SRAM

generation · layer count · projected
Each generation adds layers on the same process; density scales with layer count. Figures are projected from layout simulation and are relative to 2 nm HD SRAM at 1×. Bandwidth scales on the same roadmap, since connections run through every layer.
Technology and approach

Three ideas, one memory.

Past attempts at a new memory stalled on exotic materials or physics that scaled poorly. Z2RAM rests on three ideas that have each been proven separately, combined for the first time in a cell that meets the needs of AI infrastructure.

  1. 01

    No capacitor.

    Conventional DRAM stores each bit in a capacitor that has become nearly impossible to shrink. Nagare Labs' cell stores the bit in transistors alone: it reads fast and reads without disturbing the data. It is a proven electronic system built on well-understood physics and established, scalable manufacturing, not a materials-science bet.

  2. 02

    Transistors that don't leak.

    Built from ultra-low-leakage oxide semiconductors, the same material family in every modern display, the cell holds its data far longer than silicon can, without constant refresh, even at high junction temperatures.

  3. 03

    Deposited in layers.

    Because the cell is flat, it stacks. Layers of memory are deposited one over the other and patterned together in a single step, the same idea that brought us 3D NAND and let flash memory grow for a decade.

package / interposerAI ACCELERATOR · LOGIC DIEwired through the interposer · thousands of lanesHBM BASE DIEHBM · six bonded DRAM diesNagare Labs memorysix layers · one die · the thickness of oneTODAY · HBM BESIDE THE ACCELERATOR, BONDED DIES, WIRED THROUGH AN INTERPOSERNAGARE LABS · ONE THIN MEMORY DIE, INTEGRATED ON TOP OF THE ACCELERATOR
Today, HBM sits beside the accelerator as a stack of separately bonded DRAM dies, reached through a few thousand lanes in an interposer. Nagare Labs replaces the stack with a single memory die of six deposited layers, integrated directly on top of the accelerator and connected vertically. Thinner than one HBM die, it is also a far better path for heat leaving the logic.
Standard CMOS nodesNo EUV lithographyStandard materials, tools, and chemistriesOxide semiconductors already in volume production

The path to product runs through 300 mm production-class tooling first: proving the stacked memory at array scale on the same tools and chemistries fabs already run.

From there, production scales with foundry and packaging partners rather than a factory of our own. That is the difference between a memory that can be made and a memory that can be made in volume.

Founders

Founded by the pioneers who solved the last scaling wall.

For most of their careers, Suman Datta and Shimeng Yu advised the memory industry from the outside. Suman's transistors are in nearly every processor shipping today. Shimeng's simulator is how the industry decides which memory to build. When their own lab's results crossed from promising to buildable, they got off the sidelines and founded Nagare Labs to build it themselves.

Suman Datta
Co-founder

Dr. Suman Datta

Co-inventor of Intel's High-k/Metal Gate and FinFET transistors, the breakthroughs that carried computing from 45 nm to today. Recipient of the 2026 IEEE Andrew S. Grove Award. Pettit Chair Professor, Georgia Tech.

IEEE Fellow · National Academy of Inventors · 187+ US patents · 45,900+ citations
Shimeng Yu
Co-founder

Dr. Shimeng Yu

Creator of NeuroSim, the industry standard for modeling emerging memory and AI hardware, used by foundries and chip designers worldwide. Dean's Professor, Georgia Tech; PhD, Stanford.

IEEE Fellow · SRC Technical Excellence Award · Intel Outstanding Researcher · 40,000+ citations
Seven years in the making
2019

Suman Datta's IEEE Micro perspective proposes back-end-of-line oxide-semiconductor transistors for monolithic 3D integration.

2020

First amorphous-oxide two-transistor gain cell demonstrated at IEDM.

2024

Core material and device structure selected; device reliability optimized (VLSI).

2025

Memory bit cell experimentally demonstrated (IEDM); four-layer stacked transistor structure fabricated. Shimeng Yu's group publishes the thermal and power-delivery analysis of HBM stacked on logic that quantifies stacked DRAM's ceiling (IEEE JxCDC).

2026

Long-term stability shown (VLSI). Nagare Labs founded, with a clean institutional license from Georgia Tech and backing from Playground Global.

Board and advisors

Backed by the leaders who built modern semiconductors.

Guided by former leaders from Intel, Intel Foundry, and TSMC, and backed by Playground Global.

Board
Pat Gelsinger
Board Director · General Partner, Playground Global

Pat Gelsinger

Former CEO of Intel and VMware. Intel's first CTO, architect of the 80486, and a semiconductor industry leader for four decades.

Advisors
Sanjay Natarajan
Advisor

Sanjay Natarajan

Former Senior Vice President and General Manager of Intel Foundry Technology Research and Development, responsible for Intel's leading-edge process technology.

H.-S. Philip Wong
Advisor

H.-S. Philip Wong

Professor at Stanford University and former Vice President and Chief Scientist of TSMC. Member of the National Academy of Engineering and a founding figure in emerging memory.

Work with us

Build the next memory tier with us.

Open roles · Join the founding team

Every early hire shapes the architecture, the process, and the company.

RoleWhat you'll ownBackground we're looking for
Device & Process Integration EngineerOxide-semiconductor transistor and 3D module development on 300 mm toolingThin-film or BEOL integration; DRAM or 3D NAND process background
Memory Circuit DesignerArray, sense, and peripheral design for stacked memoryDRAM or eDRAM core design; tape-out experience

Atlanta, GA, with flexibility for the right candidates. Write to hello@nagarelabs.ai.