How Linux Dirty Cred Exploits Bypass Kernel Mitigations
Learn how Dirty Cred exploits swap file credentials in the Linux SLUB allocator to escalate privileges, bypass KASLR, and defeat modern kernel mitigations.
Modern Linux kernel hardening focuses heavily on preserving control-flow integrity. Technologies such as Supervisor Mode Execution Prevention (SMEP), Supervisor Mode Access Prevention (SMAP), Kernel Address Space Layout Randomization (KASLR), and fine-grained forward-edge Control Flow Integrity (kCFI) assume an attacker must hijack the instruction pointer to achieve privilege escalation. When attackers attempt to corrupt return addresses on the kernel stack or manipulate indirect function pointers, these mitigations trigger immediate kernel panics or restrict branch targets to authorized call sites.
Dirty Cred exploits fundamentally break this security model. Instead of altering execution paths or assembling return-oriented programming (ROP) chains, Dirty Cred targets security-critical data structures residing within kernel heap memory. By manipulating the lifecycle of kernel objects such as struct cred and struct file, an unprivileged user process can overwrite root-owned binaries or acquire full UID 0 privileges without executing a single instruction outside legitimate kernel routines.
Understanding how Dirty Cred exploits operate requires analyzing how the Linux SLUB allocator manages object caches, how memory reclamation returns physical pages to the buddy allocator, and how the virtual filesystem (VFS) decouples capability verification from subsequent write operations.
Anatomy of Kernel Credential Structures in the SLUB Allocator
The Linux kernel tracks process privileges and open file states through discrete, heap-allocated C structures. Privilege identity is encapsulated in struct cred, defined in include/linux/cred.h:
struct cred {
atomic_long_t usage;
kuid_t uid;
kgid_t gid;
kuid_t suid;
kgid_t sgid;
kuid_t euid;
kgid_t egid;
kuid_t fsuid;
kgid_t fsgid;
unsigned securebits;
kernel_cap_t cap_inheritable;
kernel_cap_t cap_permitted;
kernel_cap_t cap_effective;
kernel_cap_t cap_bset;
kernel_cap_t cap_ambient;
struct user_namespace *user_ns;
struct ucounts *ucounts;
struct group_info *group_info;
union {
struct rcu_head rcu;
call_rcu_func_t put_addr;
};
};
When a process calls setuid() or forks a worker thread, the kernel manages the lifecycle of struct cred via reference counting on atomic_long_t usage. When this counter decrements to zero, put_cred() triggers an RCU callback that releases the memory back to the allocator. Similarly, file access permissions are tracked inside struct file, defined in include/linux/fs.h:
struct file {
union {
struct llist_node fu_llist;
struct rcu_head fu_rcuhead;
} f_u;
struct path f_path;
struct inode *f_inode;
const struct file_operations *f_op;
spinlock_t f_lock;
atomic_long_t f_count;
unsigned int f_flags;
fmode_t f_mode;
struct mutex f_pos_lock;
loff_t f_pos;
struct fown_struct f_owner;
const struct cred *f_cred;
/* ... additional fields ... */
};
The SLUB allocator manages both structures in dedicated slab caches rather than generic kmalloc pools. The kernel initializes cred_jar to store struct cred instances and filp to store struct file instances. Within each slab cache, memory is partitioned into slabs, which consist of one or more contiguous physical pages assigned by the buddy allocator. A single slab contains a freelist of fixed-size slots. When a kernel subsystem calls kmem_cache_alloc(), the allocator fetches the head of the per-CPU freelist.
Crucially, the identity of a thread or file is determined entirely by the raw bit values inside these structures. If an attacker can force the kernel to point an unprivileged task's current->cred pointer to an existing root credential, or swap the struct file pointer of a writable file descriptor to point to an open read-only file like /etc/passwd, the operating system grants elevated access without questioning the legitimacy of the transition.
Why Dirty Cred Exploits Invalidate Control-Flow Integrity and KASLR
Traditional heap exploitation against the Linux kernel follows a predictable sequence: corrupt an object's function pointer table (ops), redirect execution to a gadget, pivot the stack, and execute shellcode or disable SMEP and SMAP via CR4 register alteration. Because of this, defensive architectures spent a decade implementing hardware-enforced branch verification and memory address hiding.
Dirty Cred exploits render these defenses irrelevant because they are strictly data-only attacks. No function pointers are overwritten, no executable memory is mapped, and the instruction pointer (%rip on x86_64 or PC on ARM64) never deviates from legitimate kernel text addresses.
Consider how KASLR behaves under this threat model. KASLR randomizes the base address of the kernel image and physical memory maps to prevent attackers from predicting the locations of ROP gadgets. A Dirty Cred attack does not need to know where kernel text lives. The exploit operates entirely on the relative lifecycle and placement of heap objects. If an attacker triggers a double-free or use-after-free (UAF) bug, they only need to groom the heap so that a privileged object reoccupies the memory slot previously vacated by an unprivileged object.
Similarly, kCFI verifies indirect function calls against known compile-time function prototypes. Because Dirty Cred invokes system calls through normal paths—such as standard reads, writes, and process fork routines—kCFI observes zero anomalous call sites.
Data access checks governed by SMAP prevent the kernel from dereferencing user-space memory pointers directly. In a Dirty Cred exploit, every dereferenced pointer belongs to valid kernel virtual addresses mapped within the direct physical map (page_offset_base). When an application manages thousands of file handles across worker threads using event multiplexers, the kernel maintains long-lived heap structures. As explored in How Linux epoll Works: Red-Black Trees and Ready Lists, file references remain registered inside kernel red-black trees and ready lists, holding underlying struct file allocations open across extended execution cycles. Dirty Cred takes advantage of these long-lived allocations to keep target objects parked in heap memory until the swap is executed.
The Heap-Grooming Mechanics of Slab Cache Reallocation
The central challenge in a Dirty Cred exploit involves cross-cache exploitation. A vulnerability often originates in an object managed by a generic cache, such as kmalloc-512 or kmalloc-1024, while the target credentials reside in cred_jar or filp. If slab merging (CONFIG_SLAB_MERGE_DEFAULT) is disabled, the SLUB allocator will not allocate a struct cred from a generic slab cache.
Attackers overcome this boundary by driving the source slab cache to complete page exhaustion, forcing the physical page frame to return to the kernel buddy allocator:
[ Active Slab in kmalloc-512 ]
+-------------------------------------------------------+
| Slot 0 (Free) | Slot 1 (Vulnerable) | Slot 2 (Free) |
+-------------------------------------------------------+
|
v (Trigger UAF / Free all slots)
[ Page Freed to Buddy Allocator ]
+-------------------------------------------------------+
| Order-0 Page Frame (4096 Bytes) returned to Free List |
+-------------------------------------------------------+
|
v (Spray target objects: filp / cred_jar)
[ Reallocated Slab in filp Cache ]
+-------------------------------------------------------+
| struct file A | struct file B (Occupies old Slot 1) |
+-------------------------------------------------------+
The attack follows four distinct phases:
- Heap Grooming and Slab Defragmentation: The attacker allocates thousands of objects in the vulnerable cache to fill existing partial slabs. This forces the SLUB allocator to request fresh, empty pages from the buddy allocator.
- Triggering the Free Primitive: The vulnerable object is allocated within this clean slab page. The exploit then triggers the bug to generate a dangling pointer to the target slot.
- Slab Reclamation: The attacker frees every other object residing in the same slab page. When the active object count drops to zero, the SLUB allocator calls
discard_slab(), invokingfree_the_page(). The 4 KiB physical page returns to the buddy allocator's per-CPU page set. - Target Cache Spraying: The attacker immediately sprays allocations for the target structure (such as opening thousands of read-only files to allocate
struct fileinstances, or spawning user namespaces to allocatestruct credinstances). The target slab cache requests a new page from the buddy allocator. The buddy allocator reissues the recently freed page frame.
Because the physical page has been reassigned to the target cache, the dangling pointer from the vulnerable source object now aliases directly into the memory offset of a newly allocated struct file or struct cred.
The predictability of this reclamation cycle depends heavily on system memory pressure and fragmentation. Under high allocation churn, page table management and compaction routines alter how pages are split and merged. As analyzed in Why Transparent Huge Pages on a VPS Degrade Memory Latency, allocator pressure and page frame fragmentation directly influence whether the buddy allocator delivers contiguous chunks or recycled order-0 pages, shifting the timing window required for reliable cross-cache alignment.
Winning the Time-of-Check to Time-of-Use Race Condition
The primary variant of Dirty Cred converts a read-only file descriptor into an arbitrary file write primitive by exploiting a Time-of-Check to Time-of-Use (TOCTOU) race condition in the Virtual Filesystem layer.
When an application invokes write() on a file descriptor, the kernel routes the request through vfs_write(). Inside this routine, the kernel checks whether the file is opened with write permissions:
ssize_t vfs_write(struct file *file, const char __user *buf, size_t count, loff_t *pos)
{
if (!(file->f_mode & FMODE_CAN_WRITE))
return -EINVAL;
if (!(file->f_mode & FMODE_WRITE))
return -EBADF;
if (!access_ok(buf, count))
return -EFAULT;
return file->f_op->write(file, buf, count, pos);
}
The security verification relies on file->f_mode. Once file->f_mode & FMODE_WRITE evaluates to true, the kernel proceeds to the underlying filesystem write implementation (file->f_op->write or file->f_op->write_iter). Crucially, there is an execution window between the permission check and the moment the filesystem driver commits data blocks to disk:
Thread 1 (Attacker Writer) Thread 2 (Exploit Orchestrator)
-------------------------- ------------------------------
1. open("/tmp/writable", O_RDWR)
2. sys_write(fd, payload, len)
3. vfs_write() checks f_mode:
(FMODE_WRITE is valid)
4. Trigger UAF/Free on /tmp/writable struct file
5. Spray open("/etc/passwd", O_RDONLY)
6. /etc/passwd reallocates onto old struct file slot
7. file->f_op->write(...) executes
==> Writes payload into /etc/passwd!
Under normal operating conditions, the time delta between the FMODE_WRITE check and the filesystem write callback is measured in hundreds of nanoseconds. Winning this race without stabilizing the window results in premature writes or kernel null-pointer dereferences.
To extend the execution gap from nanoseconds to hundreds of milliseconds, attackers use three primary mechanisms:
Userfaultfd Memory Stalling
The attacker places the write buffer payload into an anonymous virtual memory region registered with userfaultfd. When vfs_write() calls access_ok(), the check passes because the memory address is within user boundaries. However, when the filesystem driver subsequently copies data from the user buffer using copy_from_user(), a page fault triggers. The kernel suspends the executing thread and delegates the fault resolution to the user-space handler. The thread remains frozen inside vfs_write() after the FMODE_WRITE validation has completed, allowing the exploit thread infinite time to free the writable file and replace it with /etc/passwd.
FUSE-Backed Filesystem Trarapment
If unprivileged userfaultfd is restricted via vm.unprivileged_userfaultfd = 0, attackers mount a Filesystem in Userspace (FUSE). The source write payload is mapped from a file on the FUSE mount. When copy_from_user() touches the buffer, the kernel issues a read request to the userland FUSE daemon. The daemon intentionally delays the response, holding the target write thread frozen inside the kernel pipeline.
Disk I/O Saturation and Tracing Probe Latency
When userfaultfd and FUSE are unavailable, attackers generate massive write queues across the storage subsystem to saturate the I/O request queues. System call latency increases significantly when security monitors are attached to kernel entry points. As documented in How Linux eBPF Rootkits Evade Detection in Production, instrumentation hooks, kprobes, and tracing programs attached to functions like vfs_write alter kernel execution timing, expanding transient windows that race conditions rely on.
Kernel Defenses and Slab Virtual Isolation Mitigations
Remediating Dirty Cred required the Linux kernel security community to rethink heap isolation. Because data-only attacks do not violate architectural control-flow rules, mitigations must enforce structural isolation within the allocator itself.
Dedicated Slab Caches and Credential Separation
The most direct mitigation is preventing object aliasing. Modern kernels decouple security credentials from shared allocators entirely. In Linux 6.6 and subsequent releases, patches isolate cred_jar and filp using cache-level attributes that prohibit slab merging regardless of boot command-line configurations:
/* kernel/cred.c */
void __init cred_init(void)
{
cred_jar = kmem_cache_create("cred_jar", sizeof(struct cred),
0, SLAB_HWCACHE_ALIGN | SLAB_PANIC | SLAB_ACCOUNT, NULL);
}
By ensuring that cred_jar objects can never share slabs with generic heap allocations, an attacker cannot exploit a standard kmalloc-X vulnerability to overwrite a struct cred directly.
Virtual Slab Isolation (SLAB_VIRTUAL)
While dedicated caches prevent direct sharing, cross-cache attacks still bypass this by returning pages to the buddy allocator. To counter cross-cache reclamation, Linux researchers introduced virtual slab isolation concepts.
+---------------------------------------------------------------+
| Virtual Address Space |
+---------------------------------------------------------------+
| Zone A: Generic Objects | Zone B: Security Objects |
| (kmalloc-*, skbuff, etc.) | (cred_jar, filp, etc.) |
+--------------------------------+------------------------------+
| |
v v
Buddy Allocator Free Pool Buddy Allocator Free Pool
(Tracked for Zone A only) (Tracked for Zone B only)
SLAB_VIRTUAL reserves non-overlapping virtual memory ranges for distinct object classes. Even if a physical page backing a kmalloc-512 slab is freed to the buddy allocator, its virtual mapping cannot be repurposed for a cred_jar slab. The page cannot be mapped into the virtual address range reserved for credential jars without an expensive, audited TLB shootdown and remap sequence, breaking the rapid heap spraying required for Dirty Cred.
FMODE_CAN_WRITE Immutable Inode Binding
To stop the TOCTOU file replacement attack, modern kernels enforce deeper consistency between struct file and its underlying struct inode. Rather than relying exclusively on mutable flags stored in the file struct, write operations continuously validate that the file's associated credentials match the calling context:
static inline int file_write_not_corrupted(struct file *file)
{
return file->f_cred == current_cred() && (file->f_mode & FMODE_CAN_WRITE);
}
If an attacker frees a writable file struct and re-allocates a read-only file struct over the same memory, the mismatch between the original write context and the new read-only file parameters aborts the write with an error before the driver executes.
Hardware-Assisted Tagging (Arm MTE and KASAN)
On modern Arm64 hardware supporting the Memory Tagging Extension (MTE), each 16-byte memory granule has an associated 4-bit metadata tag. Pointers into kernel space must present matching top-byte tags to dereference the allocation. When an object is freed back to the SLUB allocator, the allocator modifies the memory tag of the slot:
Initial Allocation:
Pointer: 0xF100FFFF80104000 ---> Memory Tag: 0x1 (Matches Pointer Tag 0x1)
Object Freed & Reallocated to Credential:
Pointer: 0xF100FFFF80104000 ---> Memory Tag: 0x2 (Updated by SLUB)
Access via stale pointer fails: TAG FAULT (Synchronous Kernel Panic)
If an exploit thread attempts to access the recycled slab slot using a stale dangling pointer, the CPU generates an instant synchronous memory tag fault, terminating the kernel thread before data corruption occurs.
Conclusion
Dirty Cred demonstrated a critical blind spot in contemporary operating system security. For over a decade, defense-in-depth strategies focused on the control-flow pipeline: randomizing code addresses, disallowing execution on stack and heap pages, and authenticating branch targets with CFI.
Data-only attacks bypass this entire defensive matrix by leaving the execution path untouched. By weaponizing the operational lifecycle of the SLUB allocator and exploiting the gap between capability checks and data commits in the virtual filesystem, Dirty Cred achieves complete host takeover using ordinary system operations.
Neutralizing this class of exploit requires structural changes in the kernel heap. As virtual slab isolation, dedicated credential boundaries, and hardware-enforced memory tagging become standard across enterprise Linux deployments, kernel memory management is shifting away from pure allocation efficiency toward verifiable object compartmentalization.
Measured on our own hardware: Do Transparent Huge Pages Help on a VPS? Random Access Over 1 GiB
On this virtual server, does asking the kernel for 2 MiB transparent huge pages make a dependent random read over a 1 GiB buffer faster than the same buffer on 4 KiB pages, and how much of the buffer actually gets huge pages?
We ran it. The numbers below come from a program executed on the server hosting this site on 2026-09-27 — an AMD EPYC 9354P 32-Core Processor with 8 cores visible, 31.3 GB of memory, Linux 6.8.0-139-generic.
| Metric | Value |
|---|---|
| huge page speedup factor | 0.88 |
| accesses | 50000000 |
| buffer mib | 1024 |
| checksum | 61078 |
| huge case anon huge pages kb | 120832 |
| huge case fraction backed by 2m pages | 0.115 |
| huge pages ns per access best | 308.26 |
| huge pages ns per access mean | 313.3 |
| mode | measure |
| pages 2m in buffer | 512 |
| pages 4k in buffer | 262144 |
| repeats | 5 |
| small case anon huge pages kb | 0 |
| small pages ns per access best | 271.99 |
| small pages ns per access mean | 288.2 |
This is a shared virtual server, not an isolated test rig, so treat the absolute figures as indicative and the ratio between the two cases as the finding. The full method, the machine specification, and the complete source code are on the Do Transparent Huge Pages Help on a VPS? Random Access Over 1 GiB benchmark page, so you can check the method or run it yourself.
References
- Zhenpeng Lin, Yuhang Wu, and Xinyu Xing. "DirtyCred: Escalating Privilege and Conducting Arbitrary Writes Through Linux Kernel Credential Swapping." Black Hat USA 2022. https://raw.githubusercontent.com/Markakd/DirtyCred/master/DirtyCred_BlackHat_USA_2022.pdf
- Linux Kernel Organization. "Credentials in Linux." Kernel Documentation. https://docs.kernel.org/security/credentials.html
- Linux Kernel Mailing List. "mm/slub: Implement Dedicated and Isolated Slab Caches." https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/slub.c