Amazon S3 and its cloud equivalents are the default long-term store for modern data systems because they offer effectively infinite capacity, high durability, low per-gigabyte cost, and strong aggregate bandwidth for large, parallel accesses. That design, however, encodes assumptions tied to spinning disks: each S3 request behaves like a remote hard disk access (per-request throughput under ~100 MB/s, latencies in the tens of milliseconds), which forces every ecosystem to layer workarounds - caching, batching small writes into large objects, external metadata stores, and full-object rewrites instead of in-place updates. These constraints shape core architectures such as the data-lake stack (Iceberg + Parquet + S3) and systems that repurpose S3 as the system of record.
Those hardware assumptions are now outdated. Over two decades SSD prices have fallen until the SSD/disk cost gap is about 3×, while SSDs deliver ~100 microsecond access latency, millions of IOPS, and 4 KB granularity; datacenter networks at 100+ Gbit achieve sub-100 μs intra-datacenter and sub-millisecond regional latencies. True SSD-based, disaggregated storage with datacenter networking would remove caching and batching necessities, enable in-place updates and unified metadata/data stores, and simplify system design. Amazon’s SSD offering (S3 Express One Zone) is a constrained, higher-cost, single-AZ niche with multi-millisecond latency, highlighting provider inertia rather than a technical barrier; the necessary technology exists, so the community should build the next-generation primitives instead of waiting for incumbents.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.