An illustrative fleet-level model, not a measured claim. Set your fleet and cost assumptions to see what consolidating memory-bound serving GPUs could be worth. For a measured number on your own workload, get your number.
Illustrative model — not a measured Syntropic claim. Scales per-GPU unit economics across a fleet of memory-bound serving GPUs. The big saving is needing fewer GPUs once compression removes the memory pressure. Choose how GPU cost is paid (rented all-in, or owned + power) so energy isn't double-counted. Real figures are telemetry-verified per deployment; in compute-bound workloads savings fall toward zero.
Total estimated annual savings — full fleet
$16.4M
Basis: ~$16,425 / serving GPU / year × 1,000 GPUs
$2.1M
Syntropic fee · 12.5% of proven savings
750
GPUs eliminated
525 kW
GPU power saved
75.0%
GPU power reduction
Fleet size
Total GPUs serving memory-bound inference in the deployment / data center
Per-GPU drivers
4.0×
How memory-limited vs compute. 1 = compute-bound → no memory savings
4.27×
Default 4.27× — k4/v3 · head_dim 128 · every resident byte counted. Realized fleet savings depend on workload shape.
700 W
H100-class ≈ 700 W
Cost & physical assumptions
Neocloud rental ($/GPU-hr) already includes power — energy is shown as footprint, not added.
Energy & carbon detail (full fleet)
683 kW
Facility power saved (×PUE)
6.0 GWh
Annual energy saved
2,391 t
Annual CO₂ avoided
657 W
Spillover-movement power saved — smallest term
Sensitivity — GPU power reduction vs memory-bound factor
What Syntropic delivers
More tenants per card — public claim: 10× on shared-prefix serving workloads; the shared prefix is stored once instead of once per tenant.
Fewer GPUs for the same workload — the dominant term in this model: lower rental or capex in memory-bound serving.
Shared-prefix deduplication — the measured core of the product today; it is what produces the large multi-tenant numbers.
Per-tenant keyed isolation — bytes unusable without the tenant's key.
Aligned pricing — 12.5% of savings the meter can prove. No platform fee, no subscription; when the meter proves nothing, the customer pays nothing.
How it scales: savings are computed per memory-bound serving GPU, then multiplied by fleet size. Consolidation = the smaller of compression and memory-bound factor. In "Rented" mode the dollar total is avoided rental (power already inside it); in "Owned" mode it is avoided hardware + separately-billed electricity.
Patents pending. Illustrative model — third-party cost defaults: H100-class ~700 W, rental ~$2.40–$3.00/GPU-hr (GMI/Lambda/SemiAnalysis, 2026); off-chip data movement ~10–35 pJ/bit (SemiAnalysis/Intel/Horowitz); PUE ~1.3; representative grid carbon. Reference scale: a 50,000-GPU cluster ≈ 35 MW. Continuous model; figures derived from adjustable assumptions, not validated benchmarks. Real savings are metered per deployment at /stats.