
By Jordan Vale
The chip is rarely the one tapping its foot. When a huge AI model feels slow, it is often waiting on memory — the vast library of numbers that has to be carried in, nonstop, before the processor can write the next word.
Volantis, a San Francisco semiconductor company, says it has raised an $88 million Series A to attack that wait with light. The round, announced Oct. 1, was co-led by Lachy Groom and Abstract Ventures, with John Doerr, VXI Capital, Triatomic, and Susa Ventures participating, plus angel investors Dwarkesh Patel, Naveen Rao, and Sholto Douglas. The company's own site adds that funds raised to date now total $97 million, and names Sam Altman, Jeff Dean, and Dylan Patel among backers across that total. They are not listed on the newswire's Series A roster, so this piece does not treat them as new lead investors.
Here is the jam, in kitchen terms. A large model needs two things at once: a big pantry (memory capacity, so the whole model fits) and a wide, fast hallway into the stove (memory bandwidth, so the chip is not starved). Memory built onto the chip itself, called SRAM, is a short hallway and a tiny pantry. The high-bandwidth memory stacked beside today's graphics processors is a much bigger pantry, but the hallway still caps how fast a giant model can be served. Volantis says even newer stacked-memory ideas sit on that same tradeoff. You pick size or speed. You do not get both.
Volantis says its first system, A-1, is being designed to raise capacity and bandwidth together, by nearly a hundredfold — "nearly two orders of magnitude," in the company's words — by tying many ordinary memory chips into one pool with a photonic fabric. Photonic here means the data rides light, from tiny lasers built for this job, instead of only electrical wires. Add memory, and the pooled bandwidth is supposed to grow with it. The company says that also lets it use less expensive off-chip memory, which is how it hopes to bring the cost per token down. A token is just a chunk of text the model reads or writes. Volantis says A-1 is aimed at models exceeding 20 trillion parameters, at up to 10,000 tokens per second for a single user.
People already move data across data centers with light. Volantis's argument is that the hop from chip to memory is a different problem: more than 100 times as much data, over a much shorter distance, so the fiber gadgets built for long hallways do not simply shrink onto the package. Its version uses custom micro-VCSELs — small lasers grown in a supply chain that already makes these devices in volume — rather than separate external lasers. The company says that choice also sidesteps a tighter supply of materials used in other laser types. The claim attached to those links: end-to-end energy under one picojoule per bit. A picojoule is a tiny sip of energy. "Per bit" means per smallest piece of data moved. That number is Volantis describing the design, not a lab report from a customer.
Chief executive Tapa Ghosh put the human version of the same bet this way: as software agents take on more of a company's work, how fast they finish starts to decide how fast the company can move. The newswire's illustration is a coding agent that finishes a task in two minutes instead of 30, so a person gets more tries in a day. That is a picture of the stakes, not a stopwatch on a product you can buy.
The founding team, Volantis says, comes out of NVIDIA, AMD, Broadcom, and Ayar Labs, with earlier work that includes the first CoWoS product (a way of packing chips tightly together), the first high-volume tunable VCSELs, and early systems that put silicon photonics next to the chip. First integrated inference engines are planned for customers in 2027. The round is meant to grow the engineering team and move A-1 toward those deployments.
Ink, not pencil: nothing here is shipping. "Nearly two orders of magnitude," "under one picojoule per bit," and "20 trillion parameters" are design targets in a press release. Volantis's own Sept. 29 post does not match the newswire on every illustration — it says models over 10 trillion parameters, a boost of "over an order of magnitude," and a coding agent finishing in 30 seconds instead of 30 minutes. Where the two pages disagree, this article stays with the Oct. 1 newswire and does not average them into a third number.
Monday-morning stake: the pause while an assistant drafts an email or chews on a coding task is often a memory traffic jam, not a processor that "needs a moment to think." If the hallway between chip and memory gets wider, the useful change is conversations and agent chores that finish while you are still at the desk — not a bigger slogan on the box.
Why regular people should care
Every chat window, photo tool, and office agent is quietly rationed by how fast numbers move, and by what that electricity costs. A photonic memory pool, if it works as advertised, is one path to larger models that still answer quickly, on cheaper memory, with less energy per bit moved. That is an abundance story only if the hardware shows up. Until then it is a well-funded bet that light belongs on the shortest, busiest hallway in the machine.
What's next
Volantis says it plans to deliver the first integrated engines in 2027, and that it will share more of the architecture as A-1 gets closer to customers. The tell will be a system someone outside the company can time: tokens per second, energy per bit, and whether capacity and bandwidth really rise together. Watch the gap between the newswire and the company post, too. A startup that is still choosing between "10 trillion" and "20 trillion" in public is still designing, which is allowed — and worth remembering before anyone treats the targets as a spec sheet.