The project carried out all planned work: fast C++ implementations of UNIDACT (binary and arbitrary-alphabet), the non-binary derivation, initial work on iUNIDACT and the customised variants, and a Drop-Innovation market study sizing the opportunity and route to market.
Market context and potential impact. The study frames the impact against a global datasphere projected to grow from ~175 zettabytes in 2025 to over 1,000 by 2030 while storage covers under 15% of data generated — so compression is increasingly a hard economic constraint. It also sharpens where UNIDACT's value lies: general-purpose compression is commoditised by open-source codecs (Zstandard, Brotli) and offloaded to dedicated silicon, so competing as "a better ZIP" is not the play. Instead, the opportunity concentrates in high-value niches where legacy algorithms are structurally insufficient and where UNIDACT's three differentiators — linear decoding complexity, arbitrary-alphabet support and indexed retrieval (iUNIDACT) — map onto quantifiable pain:
- Cloud data lakes and FinOps. Cloud egress fees are a ~$43 billion/year "tax"; iUNIDACT's retrieval of parts of a file without full decompression attacks that cost directly — the largest single segment in the study.
- Genomics and life sciences. Petabyte-scale sequencing data in a quaternary (A, T, G, C) alphabet is poorly served by binary codecs; the non-binary version targets integration into the CRAM / GA4GH ecosystem.
- Aerospace and deep-space telemetry. Onboard storage and transmission limits force deletion of scientific data; a low-memory, high-ratio codec addresses the ESA/CCSDS need directly.
- Edge and IoT. Linear complexity and small-file efficiency suit battery- and bandwidth-constrained sensors, with a route to per-unit semiconductor-IP royalties.
Significance. On a bottom-up US/EU basis, the study estimates the immediately obtainable market (SOM, years 1–4) at about $19 million ARR — cloud FinOps ~$10M, genomics ~$4.35M edge/IoT ~$4.5M — scaling to a serviceable market (SAM) of ~$555 million and a total addressable market (TAM) of ~$7 billion, plus a further ~$14 billion if UNIDACT is licensed as semiconductor IP. These are the study's estimates, and the SOM is contingent on iUNIDACT reaching production readiness, which requires the further research identified below.
Key needs to secure uptake. The study's central conclusion matches our own: the binding constraint is no longer the technology but the commercial and institutional framework around it. Adoption requires:
- Further R&D to mature the results already advanced: above all iUNIDACT (the gating item for the largest, cloud, segment) to production quality, plus the customised genomic and deep-space variants — already usable today via the general-alphabet version — and a hardware/FPGA path against silicon-accelerated legacy codecs.
- Demonstration and validation on real data, beyond simulated Markov sources, with GA4GH and ESA/CCSDS partners.
- Standardisation as the distribution channel — in genomics (GA4GH/CRAM) and space (CCSDS), adoption runs through standards bodies rather than software sales.
- An IP and commercialisation vehicle — UNIDACT is not patented (software is not patentable at the EPO); it would be protected via proprietary implementations and trade secrets and taken to market through a dedicated spin-off, with a segment-specific licensing model (enterprise licences; SaaS/savings-share for cloud; IP royalties for silicon; NRE for space).
- Access to markets and finance — follow-on and national/EU funding to reach production, while the timing windows the study identifies in genomics standards, aerospace procurement and European funding remain open.
Overview of the results:
- Fast, memory-efficient C implementation of binary UNIDACT.
- Derivation and C implementation of UNIDACT for arbitrary (non-binary) alphabets.
- Foundational work on iUNIDACT and customised genomic/deep-space variants (further research needed).
- Market and commercialisation study (Drop-Innovation): competitive context, four-segment segmentation, TAM/SAM/SOM sizing (~$7B / ~$555M / ~$19M ARR), technology-transfer pathways and investment thesis.