SSRF Prevention in Small PHP Applications - A Valid URL Is Not a Safe Destination
A small PHP application may need to import an image, preview a link, test a webhook endpoint, or retrieve a remote feed. The feature sounds ordinary: accept a URL, send a request, and process the response. But whose network access is being used? The request leaves from the server, not from the visitor's browser, so it may be able to reach places the visitor cannot.
This is the boundary behind server-side request forgery, or SSRF. The problem is not simply that an input “looks like a URL.” It is that untrusted influence reaches a network client without a reliable decision about where that client is allowed to connect. A syntactically valid URL can still name a loopback service, a private address, a link-local endpoint, an unexpected protocol, or a public host that redirects elsewhere.
This article develops a defensive review model for small PHP applications. It does not provide attack payloads or claim that one validation function makes arbitrary fetching safe. The more useful goal is to identify each decision between receiving a string and opening a connection.
SSRF borrows the server's position
MITRE's CWE-918 describes SSRF as a server retrieving a supplied URL or similar request without sufficiently ensuring that it goes to the expected destination. That wording matters. The server may sit behind a firewall, share a host with administrative services, have access to private DNS, or carry credentials intended for one upstream API. A request made through that server inherits some of this position.
The result is not limited to reading a public web page. The OWASP SSRF Prevention Cheat Sheet notes that server-side retrieval may involve internal or external networks, the local machine, and protocols other than HTTP. The exact impact depends on what the process can reach and what the response is allowed to influence.
Common features create this boundary without looking like security features:
- downloading an avatar or article image from a submitted URL;
- generating a title and thumbnail for a link preview;
- letting a user test a webhook destination;
- importing a calendar, feed, or document;
- asking an automation tool to retrieve a URL found in external content.
The first design question is therefore not “Which URL validator should be used?” It is “Does this feature need to connect to a destination chosen by the user at all?”
A valid URL is not an authorized destination
URL syntax and network authorization answer different questions. Syntax validation may establish that a string fits a parser's rules. It does not establish that the host is trusted, that its current addresses are globally reachable, that the port is appropriate, or that the destination will remain the same after a redirect.
PHP makes this distinction especially easy to miss. The official parse_url() documentation says directly that the function is not a URL validator. It may accept partial or invalid URLs, does not follow one established URL standard, and may interpret input differently from another parser. That last point is critical when one component checks a host but libcurl or a proxy later interprets the actual request.
FILTER_VALIDATE_URL is a validation filter, but passing it still does not apply an application's destination policy. The PHP filter documentation describes whether data passes the chosen filter; it does not promise that the resulting destination is safe for a server to contact. Treating syntactic validity as authorization is like checking that an address is written in a valid postal format and concluding that every building at that address is open to visitors.
URLs also contain more structure than a hostname. RFC 3986 separates scheme, authority, path, query, and fragment, with user information, host, and port inside the authority. Reserved characters and percent-encoding affect interpretation. A security decision should use a well-defined parser and compare parsed components, not search the original string for a trusted substring.
The safest URL is often one the user never supplies
If the application talks only to a known API, do not accept a complete URL. Accept a narrow business identifier, such as a resource ID, then construct the request from an application-owned base URL, fixed scheme, fixed port, and controlled path encoding. This removes far more ambiguity than trying to reject every dangerous form of a general URL.
Where several destinations are legitimate, an allowlist can contain exact hosts or service identifiers. Matching should happen against the parsed and normalized host under one documented policy. A suffix check such as “ends with example.com” is not equivalent to selecting the exact host api.example.com. Ports and schemes need their own allowlists as well.
OWASP separates this known-destination case from the much harder case where the product genuinely permits arbitrary public destinations. That is an important product distinction. An internal integration, a webhook tester, and an open link-preview service do not have the same destination policy. If requirements can be narrowed, narrowing them is a security control rather than an inconvenience.
When arbitrary public URLs are required
Some features cannot use a small destination allowlist. They need a more cautious pipeline, and every failure should reject the request before a connection is opened.
1. Parse once under an explicit contract
Choose the parser model used by the runtime and retrieval stack, document it, and reject malformed or ambiguous input rather than repairing it silently. Require an absolute URL, an expected host, and no user-information component. Permit only the schemes the feature needs, normally https and perhaps http when there is a documented reason. Permit only expected ports.
This check reduces input ambiguity; it does not yet authorize the destination.
2. Resolve every address family and classify every answer
A hostname is not the endpoint used by the network connection. DNS turns it into one or more IPv4 or IPv6 addresses, and those answers may change. The application needs a policy for every returned address, not only the first convenient A record.
Checking only the familiar RFC 1918 IPv4 ranges is incomplete. The current IANA IPv4 and IPv6 special-purpose registries include loopback, link-local, shared, documentation, benchmarking, unique-local, mapped, and other special blocks with different reachability properties. A maintained IP-address library and an explicit “globally reachable under this application's policy” test are safer than a hand-written prefix list copied into one controller.
There is a difficult timing issue here. If application code resolves a hostname, approves the answer, and then gives the original hostname to another component, that component may perform a new DNS lookup and receive a different address. OWASP discusses this class of DNS-pinning or rebinding problem. Closing the check-to-connect gap can require binding the connection to approved resolved addresses while preserving the intended host for HTTP and TLS verification. The exact method depends on the HTTP client, DNS resolver, proxy, and runtime, which is why a short universal PHP helper would be misleading.
3. Treat every redirect as a new request
A public URL may return a redirect to a destination that the original check would have rejected. Automatic redirect following therefore moves the security decision after validation unless each new target goes through the same scheme, port, hostname, DNS, and address checks.
libcurl leaves CURLOPT_FOLLOWLOCATION disabled by default. Keeping it disabled is the simplest policy when redirects are unnecessary. If the feature requires redirects, use a small hop limit and inspect each target before continuing. Do not assume that restricting redirect protocols also validates the redirected host.
4. Restrict protocols in the network client
An application may intend to fetch web pages while its libcurl build supports many other protocols. The curl project's CURLOPT_PROTOCOLS_STR documentation says the default is every protocol built into libcurl. An explicit protocol list makes the client's behavior match the feature's contract.
Redirects have a separate control, CURLOPT_REDIR_PROTOCOLS_STR. This is defense in depth, not a replacement for per-hop destination checks. Runtime compatibility also matters: the string-based options were added in libcurl 7.85.0, so deployed PHP and libcurl versions must be checked before relying on them.
Put a second boundary in the network
Application validation is detailed code exposed to detailed edge cases. It should not be the only barrier between a URL-fetching worker and every service on the host or local network. OWASP recommends combining application and network-layer controls.
On a small server, that may mean running URL retrieval under a dedicated service identity or isolated worker and applying outbound firewall rules, a controlled proxy, or network namespace policy. The worker should reach only the destinations and ports its job requires, while administrative interfaces, database ports, local control sockets, and metadata services remain unreachable from that context.
Egress restrictions are not automatically simple. Public destinations can use changing address ranges, DNS may be handled by a proxy, and the same host may run several services with different needs. The policy must match the actual path a connection takes. Still, an independent network boundary changes one parser mistake from unrestricted server access into a denied connection.
Limit the fetch even after approving the destination
Destination checks answer where the request may go. They do not control how expensive or dangerous the response may be. A bounded fetch should also define:
- connection and total time limits;
- a maximum redirect count, preferably zero unless required;
- a maximum response size enforced while streaming, not only after buffering;
- accepted content types and a safe parser for the expected format;
- limits on decompression and downstream image or document processing;
- whether response bodies are stored, displayed, or discarded.
Do not forward ambient credentials, cookies, or arbitrary inbound headers to a user-chosen host. libcurl documents special handling for authorization and explicit cookie headers across hosts, but it also warns that other custom headers may contain data that should not travel to a redirected server. Build the outbound request from a minimal header set.
A fetched response should not be reflected directly into an HTML page or treated as trusted application data. SSRF prevention controls the request destination; output encoding, file validation, parser hardening, and malware policy are separate boundaries.
Review the complete decision path
A useful review follows the value rather than one function call:
- Identify every place where untrusted data can influence a server-side request.
- Ask whether a business identifier or fixed destination can replace a complete URL.
- Define allowed schemes, exact destinations or public-address policy, ports, and redirect behavior.
- Ensure validation and retrieval agree on URL interpretation.
- Check all IPv4 and IPv6 resolution results and close the gap between approval and connection.
- Re-run the policy for every redirect.
- Restrict client protocols and add an independent egress boundary.
- Bound time, bytes, redirects, parsing, and stored data.
- Log the policy outcome, destination class, timing, and failure reason without recording credentials or sensitive response bodies.
Tests should include direct IP literals, IPv4 and IPv6, multiple DNS answers, disallowed ports and schemes, redirects to rejected destinations, resolver changes between checks, oversized and compressed responses, timeouts, and proxy-enabled deployments. These are tests of the application's own policy, not a claim that a finite test set covers every URL interpretation.
Conclusion
SSRF prevention begins by recognizing that a server-side fetch is an authority decision. A URL can be perfectly valid syntax and still ask the server to cross a boundary the user could not cross directly.
The strongest small-application design avoids arbitrary URLs and constructs requests to known destinations. When open public fetching is genuinely required, no single parser, filter, or denylist is enough. Parsing, address resolution, redirect handling, protocol limits, resource bounds, and network egress controls need to agree on the same policy.
That layered answer is less satisfying than a one-line validator, but it is more honest. The final question for a URL-fetching feature is not only “Is this URL valid?” It is “At the moment of every connection, what evidence says this process is allowed to go there?”
References
- MITRE CWE, “CWE-918: Server-Side Request Forgery (SSRF),” CWE 4.20, updated April 30, 2026.
- OWASP Cheat Sheet Series, “Server-Side Request Forgery Prevention Cheat Sheet”.
- T. Berners-Lee, R. Fielding, and L. Masinter, RFC 3986: Uniform Resource Identifier (URI): Generic Syntax, January 2005.
- PHP Documentation Group, “parse_url”.
- PHP Documentation Group, “filter_var”.
- curl project, “CURLOPT_PROTOCOLS_STR explained”.
- curl project, “CURLOPT_REDIR_PROTOCOLS_STR explained”.
- curl project, “CURLOPT_FOLLOWLOCATION explained”.
- IANA, “IPv4 Special-Purpose Address Space,” updated October 9, 2025.
- IANA, “IPv6 Special-Purpose Address Space,” updated October 9, 2025.
