TwinString: Preserving String Semantics with Off-Heap Data on the JVM
This program is tentative and subject to change.
Strings are the primary mechanism through which Java applications ingest external textual data, including data read from files, databases, network interfaces, and native libraries. In data-intensive applications, such data must either be materialized as heap-allocated java.lang.String objects, incurring allocation, copying, encoding, and garbage-collection costs, or accessed through low-level and unsafe foreign-memory mechanisms that require non-standard string APIs and explicit reasoning about memory management and object lifetimes. Neither option is well suited to high-volume ingestion workloads that require both efficiency and seamless integration with existing Java code. We present TwinString, an alternative representation of java.lang.String that decouples string semantics from the physical placement of its contents. A TwinString stores its data outside the regular Java heap while preserving the standard String type and behavior expected by Java programs and libraries. VM support controls this data and manages its lifetime with garbage collection, allowing foreign textual data to be exposed as ordinary strings without introducing additional custom string types. We implement TwinStrings in GraalVM Native Image and evaluate them across several workloads, including microbenchmarks, text-processing applications over real-world datasets, and data-heavy applications using JDBC and SQLite. The results show that TwinStrings significantly reduce allocation overhead while remaining compatible with the original String API, and reduce P99.9 tail latency by up to 42.2% in realistic library and JDBC workloads by alleviating heap allocation and garbage-collection pressure.