Why Android 16-KB Page Sizes Break Native Mobile Apps
Learn how 16 KB page alignment, ELF segment padding, jemalloc arenas, and mmap offset faults break native Android libraries and how to re-align them.

Consumer mobile chipsets have operated on a fixed virtual memory page size of 4096 bytes for decades. From the initial 32-bit ARM architectures to the deployment of 64-bit ARMv8-A cores, the translation table granule configured in the translation control register (TCR_EL1.TG0) remained locked at 4 KiB. Starting with Android 15 and accelerating across flagship silicon in Android 16, device manufacturers are enabling 16 KiB translation granules at the Linux kernel level.
This transition fundamentally disrupts the assumptions baked into mobile software. When an operating system switches to Android 16-KB page sizes, native shared libraries compiled under traditional toolchains stop loading. The system terminates applications on boot with memory alignment faults, dynamic linker rejections, and aborted allocations.
The rationale for this architectural change centers on translation lookaside buffer (TLB) efficiency. As consumer mobile applications ballooned in code size and memory footprint, the hardware cost of servicing TLB misses became a primary bottleneck in application startup times and frame rendering pipelines. Moving to 16 KiB pages expands TLB coverage by a factor of four without increasing physical TLB entry counts. However, achieving that performance gain requires solving every layer where user-space software assumed a 4 KiB virtual address space.
The Hardware Economics of 16-KB Memory Granules
The ARMv8 and ARMv9 architectures provide native hardware support for three translation granules: 4 KiB, 16 KiB, and 64 KiB. The translation granule defines both the base page size and the stride of the page table walk performed by the Memory Management Unit (MMU). In a 4 KiB configuration with a 48-bit virtual address space, the MMU requires up to four levels of translation tables (L0 through L3). Each level resolves 9 bits of the virtual address, with the final 12 bits representing the byte offset within the page.
Under a 16 KiB configuration, the page offset consumes 14 bits. The page table levels each resolve 11 bits. This reduction in translation levels shortens the worst-case hardware page table walk from four memory dereferences to three (or even two depending on the configured virtual address width).
4 KiB Granule (48-bit VA):
+--------+--------+--------+--------+------------+
| L0 (9b)| L1 (9b)| L2 (9b)| L3 (9b)| Offset(12b)|
+--------+--------+--------+--------+------------+
16 KiB Granule (48-bit VA):
+--------+-----------+-----------+---------------+
| L0 (2b)| L1 (11b) | L2 (11b) | Offset (14b) |
+--------+-----------+-----------+---------------+
The primary microarchitectural win is TLB reach. Modern performance cores, such as ARM Cortex-X4 or Cortex-X925, typically feature a dedicated L1 data TLB of 48 to 64 entries, backed by a unified L2 TLB of 1,024 to 2,048 entries. In a 4 KiB system, a 1,024-entry L2 TLB maps:
$$1024 \times 4096\text{ bytes} = 4\text{ MiB of address space}$$
A mobile web browser, a high-fidelity game, or a machine learning inference runtime easily maintains an active working set exceeding 300 MiB. Under 4 KiB pages, the working set causes constant TLB thrashing, forcing the hardware page table walker to issue loads against the L3 and main system DRAM.
With 16 KiB pages, that exact same 1,024-entry L2 TLB maps 16 MiB. The reach quadruples without spending silicon area or power on wider content-addressable memory (CAM) arrays. The OS avoids the latency penalties of large zeroing and compaction operations observed when working with 2 MiB transparent huge pages, an effect analyzed in Why Transparent Huge Pages on a VPS Degrade Memory Latency.
Why not adopt 64 KiB pages, as done in some enterprise enterprise Linux distros? Internal memory fragmentation. Android devices juggle dozens of background services, isolated processes, and system daemons. Giving every small anonymous mapping and thread stack a minimum allocation quantum of 64 KiB introduces severe memory bloat, quickly exhausting 8 GB and 12 GB consumer DRAM pools. The 16 KiB granule represents the optimal point on the curve between TLB reach and memory fragmentation for mobile workloads.
Why Android 16-KB Page Sizes Break the Dynamic Linker
The most immediate failure mode on a 16-KB kernel occurs before application code executes a single instruction. When the Android dynamic linker (/system/bin/linker64 in Bionic) processes an APK's native shared objects (.so), it verifies ELF segment alignment against the system page size returned by the kernel.
An Executable and Linkable Format (ELF) binary organizes its runtime view into program headers, designated as PT_LOAD segments. A standard C++ library contains at least two PT_LOAD segments: one executable segment containing the .text section, and one writable segment containing the .data and .bss sections.
Program Headers:
Type Offset VirtAddr PhysAddr
FileSiz MemSiz Flags Align
LOAD 0x0000000000000000 0x0000000000000000 0x0000000000000000
0x000000000001a428 0x000000000001a428 R E 0x1000
LOAD 0x000000000001b440 0x000000000001d440 0x000000000001d440
0x00000000000011c0 0x00000000000028b0 RW 0x1000
Notice the Align value: 0x1000 (4096 bytes). To map an ELF segment directly from the file into the virtual memory space using the mmap system call, the operating system requires that the segment's file offset (p_offset) and virtual memory address (p_vaddr) satisfy the fundamental congruence constraint:
$$(p_vaddr - p_offset) \pmod{PAGE_SIZE} = 0$$
If this congruence rule is violated, the kernel cannot map the file page directly into the target virtual page. When a binary compiled with 0x1000 alignment is loaded on a 16-KB page kernel, p_offset and p_vaddr are aligned to 4 KiB boundaries, but rarely to 16 KiB boundaries.
In the readelf example above:
- Segment 2
p_offset:0x1b440$\implies 0x1b440 \pmod{0x4000} = 0x3440$ - Segment 2
p_vaddr:0x1d440$\implies 0x1d440 \pmod{0x4000} = 0x1440$
Here, $(p_vaddr - p_offset) \pmod{16384} = (0x1440 - 0x3440) \pmod{16384} \neq 0$. The offsets are incongruent. Bionic's linker detects this mismatch during the validation phase and instantly terminates the process:
dlopen failed: "/data/app/.../libnative.so" has bad ELF magic/alignment:
ELF segment alignment (4096) is smaller than the system page size (16384)
The dynamic linker refuses to perform costly user-space copy operations to fix up non-congruent segments at runtime, as doing so would destroy copy-on-write (COW) memory sharing between processes and defeat clean page discarding.
APK Packaging and the ZipAlign Alignment Failure
Fixing the ELF segment alignment inside the compiled shared object resolves only half of the loading pipeline. In modern Android runtimes, native libraries are rarely extracted from the APK onto the disk filesystem at install time. Instead, the application manifest sets android:extractNativeLibs="false".
When this flag is enabled, Bionic maps the .so directly out of the uncompressed APK file using mmap() targeted at the underlying file descriptor and offset. This design eliminates duplicate copies of libraries on the flash storage.
However, zip archives store entries sequentially. To map a file inside a zip directly into virtual memory, the uncompressed data of that zip entry must start at a physical byte offset within the ZIP file that is a multiple of the system page size.
Historically, the Android build chain relied on the zipalign utility configured with 4-byte or 4096-byte page alignment:
# Traditional packaging alignment
zipalign -p 4 input.apk aligned_4k.apk
If an uncompressed shared object starts at byte offset 0x102000 inside the APK, it is divisible by 4096 (0x102000 / 0x1000 = 258), but not by 16384:
$$0x102000 \pmod{0x4000} = 0x2000 \neq 0$$
When the Bionic linker attempts to call mmap():
void* ptr = mmap(desired_addr, segment_len, PROT_READ | PROT_EXEC,
MAP_PRIVATE, apk_fd, zip_entry_offset);
The Linux kernel rejects the syscall immediately, returning -1 and setting errno to EINVAL. The kernel's sys_mmap entry point explicitly validates:
if (offset_in_page(offset))
return -EINVAL;
To function on 16-KB devices, the Android Gradle Plugin (AGP) and internal build scripts must pass -p 16384 (or -p 16) to zipalign. Any native dependency embedded inside an APK that fails this 16-byte file-offset alignment generates a failure before execution begins.
Custom Allocators, jemalloc, and the madvise Pitfall
Once an application successfully clears the dynamic linker and packaging barriers, it encounters internal runtime bugs caused by hardcoded memory assumptions in C/C++ source code. The most dangerous failures occur inside custom memory allocators, garbage collectors, and pool engines.
Many legacy C and C++ libraries assume that memory pages are universally 4096 bytes. Codebases frequently contain macro definitions like:
// Broken legacy assumption
#define PAGE_SIZE 4096
#define PAGE_MASK (~(PAGE_SIZE - 1))
#define ROUND_UP_PAGE(addr) (((addr) + 4095) & ~4095)
When an allocator manages memory using a hardcoded PAGE_SIZE of 4096 on a 16-KB kernel, its internal accounting diverges from the physical translation mappings established by the kernel.
Consider an allocator that tracks dirty memory chunks and releases idle pages back to the operating system using madvise() with the MADV_DONTNEED advice. The allocator intends to free a 4096-byte unused range inside a pooled 64 KiB arena:
// Allocator attempts to free a sub-page range
madvise(buffer_address + 4096, 4096, MADV_DONTNEED);
The behavior of madvise(MADV_DONTNEED) under Linux is strictly page-aligned. When presented with unaligned addresses or byte lengths that do not match the system page size, the kernel has two ways to resolve the request, neither of which the allocator anticipates:
- Rejection: The kernel returns
-EINVALbecause the target address is not aligned to the system's 16 KiB translation boundary. The memory is never returned, leading to uncontrolled memory leaks and Out-Of-Memory (OOM) kills. - Over-discarding: If the address is aligned to 16 KiB, but the length spans only a partial page, the kernel's rounding logic may discard the entire 16 KiB physical page backing the region.
If the kernel discards the entire 16 KiB page, active, live objects residing in the adjacent 12 KiB of memory within that same page frame are zeroed out instantly. The next time the application accesses those neighboring objects, it dereferences null bytes or crashes with random memory corruption.
Allocators such as modern versions of jemalloc and scudo (Android’s default allocator) avoid this by eliminating compile-time page constants. They dynamically initialize their base arena scales by invoking sysconf(_SC_PAGESIZE) or getpagesize() during runtime startup:
#include
static size_t system_page_size = 0;
void allocator_init(void) {
system_page_size = sysconf(_SC_PAGESIZE);
}
Failure to query the runtime page size also breaks custom thread stack allocation. If an application runtime provisions a thread stack with a 4 KiB guard page at the boundary using mprotect(stack_bottom, 4096, PROT_NONE), the call fails with -EINVAL. The kernel refuses to apply memory protection flags to partial pages. The guard page mapping must be scaled to 16 KiB.
File I/O Offsets, DMA Alignment, and Serialization
Beyond memory allocators, systems handling asset serialization and storage caching frequently break on 16-KB kernels. Game engines, audio engines, and database systems (such as SQLite or custom RocksDB forks) heavily use memory-mapped file I/O (mmap).
In many custom asset bundling formats, multiple resources (textures, audio streams, level meshes) are concatenated into a single archive file. The asset manager constructs an index containing the byte offset and length of each asset. To eliminate file copy overhead, the engine maps individual assets directly into user space:
Asset* map_asset(int fd, off_t asset_offset, size_t asset_len) {
// Fails on 16-KB kernels if asset_offset is only 4K aligned!
void* data = mmap(NULL, asset_len, PROT_READ, MAP_SHARED, fd, asset_offset);
if (data == MAP_FAILED) {
return NULL;
}
return (Asset*)data;
}
If an asset was packed at a 4-KiB aligned offset (e.g., offset 0x3000), the mmap call fails on a 16-KB kernel because 0x3000 % 0x4000 != 0. Software must handle this by calculating the page-aligned base offset, mapping from that aligned boundary, and applying an internal pointer adjustment:
Asset* map_asset_safe(int fd, off_t asset_offset, size_t asset_len) {
size_t page_size = sysconf(_SC_PAGESIZE);
off_t aligned_offset = asset_offset & ~(page_size - 1);
size_t offset_diff = asset_offset - aligned_offset;
size_t map_length = asset_len + offset_diff;
uint8_t* base = (uint8_t*)mmap(NULL, map_length, PROT_READ, MAP_SHARED, fd, aligned_offset);
if (base == MAP_FAILED) {
return NULL;
}
return (Asset*)(base + offset_diff);
}
This interaction between page alignment and low-level device operations mirrors direct storage constraints. DMA controllers and high-throughput hardware queues demand strict physical address continuity and page-aligned buffers. Similar structural requirements govern storage subsystems, such as those covered in How NVMe Host Memory Buffer Works in Client DRAM-Less SSDs, and the batched submission architectures analyzed in io_uring vs pread: How Batched NVMe I/O Actually Scales.
Rebuilding Native Code: Linker Flags and Binary Bloat
Resolving the 16-KB page size transition requires modifying compiler and linker invocations across all native libraries. The dynamic linker requires that ELF program headers align to 16 KiB. This is controlled via the GNU Gold or LLVM lld linker flags.
For projects utilizing CMake within the Android NDK, developers must instruct the linker to establish a maximum page size of 16384 bytes using -Wl,-z,max-page-size=16384:
# CMakeLists.txt configuration
cmake_minimum_required(VERSION 3.22)
project(EngineNative)
# Enforce 16-KB segment alignment across all target architectures
add_link_options("-Wl,-z,max-page-size=16384")
For legacy configurations relying on Android.mk:
# Android.mk configuration
LOCAL_LDFLAGS += -Wl,-z,max-page-size=16384
Starting with NDK r27, the toolchain applies -Wl,-z,max-page-size=16384 by default for 64-bit ARM architectures (arm64-v8a). However, older NDK releases (NDK r26 and earlier) default to 4096 bytes.
The Problem of Binary Bloat
Altering the linker page size configuration introduces a trade-off: binary size bloat. When the linker enforces a 16-KB alignment boundary between the executable code segment (RX) and the read-write data segment (RW), it must insert padding bytes into the ELF file on disk so that the physical file offsets mirror the required virtual memory alignment.
4-KB Aligned ELF:
[ .text segment ] [ Pad: up to 4095 bytes ] [ .data segment ]
16-KB Aligned ELF:
[ .text segment ] [ Pad: up to 16383 bytes ] [ .data segment ]
On average, setting -z max-page-size=16384 adds between 8 KiB and 12 KiB of zero-padding per shared library. In modular applications containing dozens or hundreds of granular .so files, this padding can add several megabytes to the download size.
To minimize disk bloat while retaining 16-KB runtime compatibility, developers configure the common page size alongside the maximum page size:
-Wl,-z,max-page-size=16384 -Wl,-z,common-page-size=4096
This configuration instructs the linker to lay out the virtual address space on 16-KB boundaries while permitting the file layout to use 4-KB steps where permitted. However, Bionic's ELF loader strictly enforces $(p_vaddr - p_offset) \pmod{PAGE_SIZE} = 0$. If p_offset and p_vaddr are not congruent to 16 KiB, the binary will still fail on a 16-KB kernel. True compatibility requires full 16-KB file alignment padding.
Verifying Shared Libraries
To verify whether a shared object complies with 16-KB kernels, inspect the program headers using llvm-readelf or readelf:
$ aarch64-linux-android-readelf -l libtarget.so
Examine the output for the LOAD entries:
Program Headers:
Type Offset VirtAddr PhysAddr
FileSiz MemSiz Flags Align
LOAD 0x0000000000000000 0x0000000000000000 0x0000000000000000
0x0000000000042a10 0x0000000000042a10 R E 0x4000
LOAD 0x0000000000044000 0x0000000000048000 0x0000000000048000
0x00000000000021b0 0x0000000000003000 RW 0x4000
Two conditions confirm full compliance:
- Every
LOADsegment has anAlignvalue of0x4000(16384 in hexadecimal) or greater. - For every
LOADsegment, $(VirtAddr - Offset) \pmod{0x4000} == 0$. In the second segment above:(0x48000 - 0x44000) = 0x4000, which leaves a remainder of 0 when divided by0x4000.
If Align shows 0x1000, the library will fail to load on any device running a 16-KB kernel.
Conclusion
The transition to 16-KB page sizes represents a significant architectural evolution for the consumer Android ecosystem. The microarchitectural justification is sound: expanding hardware TLB reach by 4x eliminates a pervasive memory translation bottleneck, directly improving application launch performance, background switching overhead, and power efficiency across high-density workloads.
Yet, this shift exposes assumptions deeply embedded across the native software stack. Resolving these issues requires an end-to-end audit:
- Recompiling all native code using
-Wl,-z,max-page-size=16384to ensure ELFPT_LOADsegments align with the hardware page granule. - Updating APK packaging pipelines to ensure uncompressed
.soassets are aligned to 16-byte boundaries viazipalign -p 16. - Eliminating static
PAGE_SIZEmacros in C and C++ source code, replacing them with dynamic runtime discovery viasysconf(_SC_PAGESIZE). - Re-architecting custom memory allocators and file mapping logic to prevent destructive
madvise(MADV_DONTNEED)discards and misalignedmmapinvocations.
As devices running 16-KB kernels reach consumer hands, teams that update their build chains and remove 4-KB assumptions will capture performance benefits without stability regressions.
Measured on our own hardware: Do Transparent Huge Pages Help on a VPS? Random Access Over 1 GiB
On this virtual server, does asking the kernel for 2 MiB transparent huge pages make a dependent random read over a 1 GiB buffer faster than the same buffer on 4 KiB pages, and how much of the buffer actually gets huge pages?
We ran it. The numbers below come from a program executed on the server hosting this site on 2026-09-27 — an AMD EPYC 9354P 32-Core Processor with 8 cores visible, 31.3 GB of memory, Linux 6.8.0-139-generic.
| Metric | Value |
|---|---|
| huge page speedup factor | 0.88 |
| accesses | 50000000 |
| buffer mib | 1024 |
| checksum | 61078 |
| huge case anon huge pages kb | 120832 |
| huge case fraction backed by 2m pages | 0.115 |
| huge pages ns per access best | 308.26 |
| huge pages ns per access mean | 313.3 |
| mode | measure |
| pages 2m in buffer | 512 |
| pages 4k in buffer | 262144 |
| repeats | 5 |
| small case anon huge pages kb | 0 |
| small pages ns per access best | 271.99 |
| small pages ns per access mean | 288.2 |
This is a shared virtual server, not an isolated test rig, so treat the absolute figures as indicative and the ratio between the two cases as the finding. The full method, the machine specification, and the complete source code are on the Do Transparent Huge Pages Help on a VPS? Random Access Over 1 GiB benchmark page, so you can check the method or run it yourself.
Measured on our own hardware: io_uring vs pread: Random Reads at Queue Depth 32
How many random 4 KiB reads per second can one thread sustain with serial pread() syscalls, compared with io_uring submitting 32 at a time, when the page cache is taken out of the picture?
We ran it. The numbers below come from a program executed on the server hosting this site on 2026-09-27 — an AMD EPYC 9354P 32-Core Processor with 8 cores visible, 31.3 GB of memory, Linux 6.8.0-139-generic.
| Metric | Value |
|---|---|
| io uring speedup factor | 7.4 |
| block bytes | 4096 |
| blocks read | 20000 |
| io uring best seconds | 0.6754 |
| io uring iops | 29614 |
| o direct | true |
| pread best seconds | 4.9975 |
| pread iops | 4002 |
| pread mean latency us | 249.87 |
| queue depth | 32 |
| repeats | 3 |
This is a shared virtual server, not an isolated test rig, so treat the absolute figures as indicative and the ratio between the two cases as the finding. The full method, the machine specification, and the complete source code are on the io_uring vs pread: Random Reads at Queue Depth 32 benchmark page, so you can check the method or run it yourself.
References
- Android Open Source Project. Support 16 KB Page Sizes. Google Developer Documentation, 2024. https://developer.android.com/guide/practices/page-sizes
- Ulrich Drepper. How To Write Shared Libraries. Red Hat, Inc., 2011. https://www.akkadia.org/drepper/dsohowto.pdf
- Arm Limited. Arm Architecture Reference Manual Armv9, for Armv9-A architecture profile. Arm Developer Documentation, 2024. https://developer.arm.com/documentation/ddi0487/latest