DNA Data Storage: The First Real Step Toward Molecular IT

We are producing data faster than we can responsibly preserve it. Estimates put the amount created, captured and copied each year at around 180 zettabytes, a number so large it stops meaning anything until you ask how much of it will still exist a year later. Most is transient: duplicated, transmitted, processed and discarded. IDC has estimated that about 2% is retained, but 2% of 180 zettabytes is still roughly four zettabytes of photographs, health records, satellite imagery, AI models, sensor logs, video and scientific measurements that someone has decided must survive. Written to today’s highest-capacity LTO-10 tapes, that annual remainder alone would consume about 100 million cartridges at their native capacity, or 40 million only if the data compressed perfectly. To record it within a year would require more than 317,000 drives streaming without interruption, before allowing a single drive for redundancy, retrieval or migration. The problem is no longer simply how much data we create. It is how we keep even the small fraction that matters readable for decades.

For seventy years, that home has mostly been magnetic media. Tape, in particular, remains the backbone of long term cold archival data storage even when it resides in the cloud; the data you rarely touch but can never afford to lose. It's cheap, it's mature, and it keeps improving. But tape is also physical in the least convenient way: it degrades, it requires climate-controlled facilities, it needs robotic libraries the size of shipping containers, and a single year of global data would fill roughly tens of thousands of the largest libraries available. Formats become obsolete. Drives wear out. Every few years, someone has to migrate petabytes from one generation of hardware to the next, at real cost and real risk.

This is the backdrop against which DNA data storage has quietly moved from a laboratory curiosity to something research institutions, standards bodies, and even national libraries are now taking seriously — as demonstrated by the recent SCDNA26 conference in Rome, where the DNA Data Storage Alliance, the EIC's DigNA portfolio, and organizations like the U.S. Library of Congress compared notes on how close the technology actually is to production use.

Why DNA, and why now

The pitch for DNA as a storage medium is almost absurdly simple: life has already solved the problem we're struggling with. Evolution needed a molecule that could encode enormous amounts of information, survive harsh conditions, and remain interpretable across immense timescales — and it produced DNA. Repurposing that molecule for digital archives inherits three properties that no synthetic storage medium can currently match.

Density. At its theoretical limit, DNA can store on the order of a billion gigabytes per cubic millimeter. A shipping container's worth of tape could, in principle, be replaced by something the size of a few sesame seeds.

Durability. DNA can last for millennia. The oldest DNA sample scientists have successfully sequenced is roughly two million years old. In more direct tests, Microsoft researchers exposed DNA-encoded files to neutron radiation equivalent to 4.4 million years sitting in New York City — the files came back intact. Tape and hard drives don't survive their own warranty periods with that kind of confidence. Nowadays we can certainly industrialize DNA-based storage solutions to last many decades if not centuries.  

Permanence of the reading tool. Every other digital format eventually faces the "floppy disk problem" — if the media and its data survive, no device can read it anymore. DNA doesn't have this problem. Because DNA sequencing is fundamental to medicine and biology, humanity has an enormous, self-renewing incentive to keep building better tools to read it. As long as there is life science, there will be DNA readers.

Put those three properties together and you get something genuinely new in the history of data storage: a medium whose reliability increases, rather than decreases.

Is the real bottleneck today writing or reading DNA?

None of this makes DNA storage a drop-in replacement for your SSD, and it isn't trying to be. The honest, near-term use case — echoed independently by researchers in both Rome and Boston — is cold storage: the archival tier for data that is too important to delete but too rarely accessed to justify keeping on spinning disks or even on tape. Think regulatory records, scientific archives, medical history, cultural heritage, or the "long tail" of enterprise data that compliance requires you to retain for decades. Furthermore, strengthening data protection and possibly data security against rising AI-powered threats becomes a priority for valuable digital assets.

The obstacle standing between DNA storage and that future is synthesis — the actual production of DNA strands. Writing DNA one base at a time is slow and relatively expensive, and it hasn't yet found its industrial-scale answer. While several routes are being evaluated, while at Catalog our engineers took a different route: rather than maximizing density, our "Shannon" machine prints DNA the way an inkjet printer prints ink, trading some data density for dramatically faster and cheaper writing.

Meanwhile, coding efforts are attacking other major challenges, designing encoding schemes that reduce the number of synthesis cycles needed in the first place — treating the code itself as a lever for cost and speed, not just error correction.

None of these approaches is a finished product yet. But three independent lines of attack — parallel chemistry, alternative print-style synthesis, and smarter coding — converging on the same bottleneck is exactly the pattern you'd expect to see shortly before a technology crosses from lab to industry.

Given these factors, one could easily conclude that writing DNA represents the primary bottleneck. While writing remains an essential obstacle to tackle, reading DNA will likely emerge as the next critical hurdle. Overcoming it is crucial to expanding capacity beyond genomic-scale datasets, enabling faster and higher-quality quality control, and executing more sophisticated molecular data operations.

Ultimately, both reading and writing DNA present substantial challenges that must be overcome before DNA can gain widespread adoption as a dependable storage medium.

Beyond storage: the case for DNA compute

The more interesting long-term story may not be storage at all — it's computation. Once information exists as molecules rather than as electrical charge on silicon, an entirely different category of operation becomes possible: chemical computation performed directly on the data, without ever converting it back into bits.

Our engineers have already demonstrated simple search — finding a specific sequence within a much larger pool of DNA — using chemistry alone. Extend that idea and you can imagine comparing databases, matching patterns, or filtering large datasets with a fraction of the energy a silicon supercomputer would require for the equivalent task. Cache DNA, a spinout from MIT's Bathe Lab, has pushed this further with molecular "barcoding": DNA archives tagged with short identifying sequences that fluoresce under specific chemical probes, allowing specific files — images tagged "cat," "orange," "domestic," for instance — to be located and verified without sequencing the entire archive. The same underlying idea has been floated for public health surveillance: tagging viral RNA samples with metadata (location, date, patient age band) that could later be searched chemically to trace how a pathogen spread.

This is the outline of what researchers are starting to call Molecular IT — an information stack where storage, indexing, retrieval, and even parts of computation happen in the same chemical substrate, rather than being split between magnetic media and silicon processors. It's a genuinely different architecture, not just a denser hard drive.

What this means for organizations planning for the long term

DNA storage will not replace the cloud, and no one seriously building the technology claims otherwise. Data centers, SSDs, and tape aren't going anywhere in the next decade. But for any organization that has to guarantee data integrity across timescales measured in decades rather than fiscal quarters — libraries, government archives, healthcare systems, scientific institutions, and increasingly any company sitting on regulated "must-retain" data — DNA storage is worth watching closely, or worth piloting now.

The technology still has to prove itself as boring, dependable IT infrastructure: interoperable formats, predictable costs, verifiable integrity, and compatibility with the systems that already run the world's data. That is exactly the work standards bodies like the DNA Data Storage Alliance are doing today. But the direction of travel is clear. We are watching the early stages of a shift from electronic memory to molecular memory — and, potentially, from electronic computation to molecular computation alongside it. The organizations that start understanding this technology now will be the ones best positioned to use it once the economics catches up with the chemistry.