How XML External Entity Attacks Work and Where They Fail
XXE exploits parsers that trust external data sources, allowing attackers to read local files or trigger server-side requests without executing code directly.

XML External Entity attacks exploit parsers that process external entities within XML documents. An attacker injects a malicious entity definition pointing to a local file or internal server address. The parser retrieves this data and embeds it in the response, leaking sensitive information or enabling server-side request forgery.
Anatomy of the Parser Vulnerability
The root cause of an XML External Entity (XXE) attack lies in how an XML parser processes document structure. XML is a markup language designed to store and transport data. It allows documents to define entities, which are variables that hold data. An external entity is a specific type that retrieves content from a Uniform Resource Locator (URL) or a local file path.
When a parser encounters an external entity reference, it fetches the content. This feature was designed for document modularity, allowing large documents to reference external stylesheets or data fragments. However, if the parser processes these references automatically, it creates a bridge between the application and the underlying file system or network. The application treats the retrieved data as part of the valid XML stream.
This behaviour is often enabled by default in older libraries or misconfigured modern ones. The vulnerability exists because the parser distinguishes little between a harmless external stylesheet and a malicious file read command. You must understand that the flaw is not in the XML syntax itself, but in the library’s willingness to resolve external references.
Stage 1: Injection of Malicious Payload
The attack begins when an application accepts XML input from a user. This input might come from an API endpoint, a file upload feature, or a web form. The attacker crafts an XML document that includes a custom entity definition. This definition points to a resource the attacker wishes to access, such as /etc/passwd on a Unix system or C:\Windows\win.ini on Windows.
Imagine an application that accepts XML for user profile updates. The attacker submits a payload that defines an entity named xxe which points to a local configuration file. The parser sees this definition and prepares to resolve it. The injection does not require authentication if the endpoint is open, though authenticated access expands the scope of targetable files.
The critical element here is the DOCTYPE declaration. This section of the XML document defines the document type and allows entity declarations. Without a DOCTYPE, you cannot define external entities in standard XML. The attacker relies on the application passing this entire structure to the parser without sanitisation.
Stage 1: Crafting the Entity
- Identify an input field that accepts XML or converts user input to XML.
- Construct a DOCTYPE declaration with an internal subset.
- Define an external entity that references a target resource.
- Reference this entity in the XML body to trigger resolution.
Stage 2: Parser Resolution and Data Retrieval
Once the parser receives the malformed XML, it begins processing the document tree. It encounters the entity reference and initiates a request to the specified location. This step occurs on the server side, using the privileges of the application process. If the application runs with high privileges, the attacker can access system files that are otherwise inaccessible.
The parser fetches the content of the target file. It does not execute the file; it simply reads the bytes. This makes XXE particularly dangerous for configuration files, credential stores, or private keys. The data is loaded into memory as part of the XML document object model.
If the target is a remote URL, the parser performs a server-side request forgery. This allows the attacker to probe internal networks that are not exposed to the internet. The server acts as a proxy, sending requests to internal services and returning the responses. This bypasses firewall rules that restrict inbound traffic from external IPs.
Stage 3: Data Exfiltration via Response
The final stage involves returning the retrieved data to the attacker. In a reflective XXE attack, the application embeds the entity’s content in the HTTP response. The attacker sees the file contents directly in the browser or HTTP client. This is the simplest form but relies on the application echoing the parsed data back to the user.
Many modern applications do not reflect XML input directly. They process it and return a JSON response or a transformed view. In these cases, the attacker cannot see the data immediately. This leads to blind XXE attacks. The attacker must use side channels to exfiltrate data. They might define an entity that sends the retrieved file contents to a web server controlled by the attacker.
Imagine an application that processes an XML payment request. The attacker defines an entity that triggers an HTTP request to an external server, carrying the contents of a database configuration file in the URL parameters. The application logs an error or continues processing, but the data has left the environment. This method is stealthier and harder to detect without monitoring outbound traffic.
| Stage | What happens | Where it can be stopped |
|---|---|---|
| Injection | Attacker submits XML with external entity definition | Input validation and schema enforcement |
| Resolution | Parser fetches content from local or remote source | Disabling external entity support in parser |
| Exfiltration | Data is returned in response or sent via side channel | Network egress filtering and response sanitisation |
Mitigation Strategies and Code Review
Preventing XXE requires changes at multiple layers. The most effective measure is to disable external entity processing in the XML parser. Most modern parsers have a flag or setting to turn this feature off. You should verify that this setting is applied globally, not just for specific endpoints. Relying on individual developers to remember this setting is a common failure point.
Input validation is a secondary defence. You can restrict XML input to a specific schema using Document Type Definition (DTD) or XML Schema Definition (XSD). This limits the structure of the input, reducing the attack surface. However, schema validation does not always prevent entity resolution if the parser is misconfigured. It is a belt-and-suspenders approach.
Secure code review should focus on how XML libraries are instantiated. Check for deprecated parsers that do not support disabling external entities. If you are using legacy software, consider replacing the library rather than patching the configuration. Dependency updates can introduce newer, safer versions of parsers, but you must ensure the update includes the security fix.
See also: How to Prevent Hard-Coded Credentials in Source Code · Patch Management: Eight Questions Answered for Stability
The Role of Configuration and Maintenance
Even with code changes, configuration drift can reintroduce vulnerabilities. Applications deployed across multiple environments may have different parser settings. A secure setting in development might be overridden in production for compatibility reasons. You must standardise parser configurations across all environments.
Patch management plays a role here. Vendors often release updates that tighten default parser settings. Applying these updates ensures you are not relying on manual hardening that might be forgotten during redeployment. However, patches do not fix bad coding practices. If developers explicitly enable external entities in their code, a patch will not help.
End-of-life software is a significant risk. Older versions of XML libraries may not support disabling external entities at all. You cannot secure a parser that lacks the feature. Upgrading such libraries is mandatory, not optional. This aligns with broader software composition analysis practices, where you track the security posture of all third-party components.
Blind XXE and Network Detection
Blind XXE attacks are harder to detect because they do not reflect data in the response. They rely on out-of-band techniques. The attacker monitors their own server for incoming requests. This requires the victim server to have outbound internet access. If the server is air-gapped, this method fails.
Network monitoring tools can detect anomalous outbound XML-related traffic. Look for XML parsing errors or unexpected HTTP requests to external domains. These indicators suggest a blind XXE attempt. However, this is a reactive measure. It alerts you after the attack has occurred, not before.
Mitigations and workarounds for blind XXE include restricting outbound network access from application servers. If the server cannot reach the internet, it cannot exfiltrate data via HTTP or DNS. This is a strong defence but may impact legitimate functionality. You must balance security with operational needs.

Limitations of Common Defences
Web application firewalls can filter known XXE payloads. They look for patterns like <!ENTITY or SYSTEM. However, attackers can obfuscate these patterns using character encoding or whitespace. A firewall might miss a variant that uses UTF-16 encoding or splits the entity declaration across multiple lines.
Input length limits are ineffective. An XXE payload can be very short. The damage depends on the parser’s behaviour, not the size of the input. You cannot mitigate XXE by limiting the amount of data a user can send.
Authentication does not prevent XXE. If an authenticated user can submit XML, they can trigger the vulnerability. The risk increases because authenticated users often have access to more sensitive resources. Do not assume that login walls protect against this class of attack.
Key takeaways
- Disabling external entity processing in the parser configuration is the most effective defence against XXE.
- Blind XXE variants exfiltrate data via side channels, making detection difficult without network monitoring.
- Input validation alone is insufficient because the vulnerability lies in the parser’s handling of the XML structure, not the content itself.
XML External Entity attacks exploit the parser’s trust in external data sources, allowing file reads and network probes. Disable external entity processing in your XML parser configuration to eliminate this risk.
Frequently asked questions
Can I prevent XXE by using JSON instead of XML?
Yes, JSON does not support external entities or DTDs. Switching to JSON for data interchange eliminates the XXE attack surface entirely.
Does sanitising the XML input fix XXE?
Sanitisation can remove known malicious patterns, but it is not reliable. Attackers can bypass filters using encoding tricks. Disabling entity processing in the parser is the only sure fix.
How do I know if my parser is vulnerable?
Check the documentation for your XML library. If it allows external entity resolution by default, it is vulnerable. You must explicitly disable this feature in your code.
Can XXE lead to remote code execution?
XXE itself does not execute code. However, it can be used to retrieve credentials or configuration files that enable further attacks, including remote code execution.
How this guide was produced: written by the Malware Brief editorial team with AI assistance, checked against the public references listed below, and reviewed when the facts change. See our editorial policy or report an error.



