Skip to content

Image: portswigger.net · rights & removal

Executive Summary

A researcher collaborated with another individual to explore methods for exfiltrating larger data tokens using CSS injection, moving beyond previous demonstrations. The research focused on reconstructing long tokens by stitching together smaller chunks extracted from a web link. The methodology involved analyzing the overlaps between short text fragments within the source material and using graph theory to determine potential sequences of these chunks that form complete tokens. The process involved setting challenges: reproducing known results with minimal CSS and assessing token extraction length within a fixed CSS budget. Testing revealed that by intelligently grouping potential character joins, the required amount of CSS could be significantly reduced compared to previous methods.

Facts Only

* A researcher collaborated with Alex to exfiltrate larger tokens using tools like DOM Invader.
* The process involved stitching together smaller chunks to extract longer tokens without recursive loading.
* The method used graph theory to connect overlapping text fragments from a link into a complete token sequence.
* Tests demonstrated that using two-character and three-character chunks allowed for reconstruction, with certain grouping strategies yielding high success rates under specific thresholds (e.g., 90% or 99% probability).
* For a 12-character token, the shortest successful CSS required 1.33 MB, and extracting 640 characters within Gareth's original byte budget was achieved with up to five candidates.
* The final tested setup for achieving a 99% one-guess target for a 12-character token used 1.33 MB of CSS and 25,088 selectors.
* The technique was applied to generate test tokens by creating new URLs with randomized data for empirical validation.
* The process was demonstrated using Portswigger labs, where an AI tool was adapted to exploit CSS injection for token extraction across various web resources.

Full Take

The core implication of this research is a shift from brute-force methods relying on sequential loading to an overlapping, combinatorial approach that leverages relational structure (graph theory) for data reconstruction. The initial breakthrough moves beyond simple exfiltration by treating the resulting code as a constrained system where the goal is not just extraction but efficient pathfinding through potential overlaps. The pattern of increasing chunk length and testing more complex joins reveals a tension between computational efficiency (minimizing CSS size) and epistemic certainty (ensuring the reconstructed token is correct). This exploration suggests that long-range data manipulation in web contexts may be less about raw brute force and more about exploiting the inherent, often overlooked, syntactic relationships within the injected code. The final results showing reductions in required CSS by orders of magnitude suggest a systemic vulnerability where complexity can be managed through structured modeling rather than sheer volume.
What assumptions about the efficiency of data extraction versus certainty are we making when prioritizing smaller stylesheets? How does this technique change the perceived security boundary between code injection and data exfiltration? Does focusing on constructing optimal join groups risk overlooking novel, more complex semantic relationships that might offer exponential gains in token recovery?

From the original · PortSwigger Research

Researcher Published: Monday, 5 October 2026 at 15:04 UTC Updated: Tuesday, 6 October 2026 at 08:33 UTC I'm delighted to introduce Alex, my fellow swigger who I've collaborated with in the past with tools like DOM Invader. He showed me that it's possible to exfiltrate larger tokens than demonstrated in my original "CSS: the bomb inside your inbox" post.
Read the full story at portswigger.net

Sentinel — Human

Confidence

This text appears to be a genuine, detailed account of advanced technical research and experimentation, exhibiting the idiosyncratic style and deep contextual knowledge typical of expert human authorship.

Signals Detected
low severity: Erratic flow mixed with highly specific technical vocabulary; strong, personalized voice.
low severity: Deep, self-directed exposition of a complex, novel technical process with clear internal logic and sequential testing.
low severity: Direct description of experimental results, method derivation (e.g., graph theory), and explicit benchmarking against prior work by 'Gareth'.
low severity: Specific, quantifiable metrics (e.g., 1.33 MB CSS, 257.69 MB budget) and references to specific, verifiable tooling/labs (Burp AT demos).
Human Indicators
The narrative structure follows the progression of a personal research project, involving discovery, theoretical application (graph theory), experimental testing with numerical results, and practical tool integration.
The tone oscillates between technical explanation and personal enthusiasm ('I'm delighted', 'I was blown away'), characteristic of an engaged researcher sharing findings.
Smashing the token limit with overlapping fragments | Huntaegis