html-rewriter
This module is available to use in your EdgeWorkers functions to consume and rewrite HTML documents. The html-rewriter module includes a built-in parser that emulates standard HTML parsing and DOM-construction.
📘 To learn more, go to the Dynamic Content Assembly using the html-rewriter use case in this guide.
An EdgeWorkers function can register callbacks on CSS selectors. When the parser encounters an element matching the selector, it executes the callback. The callback can insert new content around the element, modify the tag attributes, or remove the element entirely.
🚧 The HtmlRewritingStream does not escape inserted text. This means that you need to validate and escape user-supplied text to prevent cross-site scripting (XSS) vulnerabilities.
HtmlRewritingStream Object
There are three steps to using the html-rewriter.
-
Create a new
HtmlRewritingStream.
You need to create a new rewriter for each stream, because it’s a stateful HTML parser. -
Add one or more handlers using the
onElement()method.
The handlers call functions on their argument to modify the stream. -
Pipe an HTML stream through the rewriter object.
Consider this EdgeWorkers function that inserts a beacon script into a webpage.
The handler passed to the onElement() method inserts a new script tag into the HTML document.
This example shows how to insert a script tag.
Here’s the updated code snippet that includes the specified script tag. A handler runs when it encounters the open tag. The operations that run on the handler’s parameter can either occur immediately, or when the element closes.
👍 You can associate multiple handlers with an
HtmlRewritingStream.
Streaming
You can use the HtmlRewritingStream in pipe chains. It acts as a transform stream that you can use with ReadableStream.pipeThrough() and ReadableStream.pipeTo().
When reading, instances of HtmlRewritingStream expect ArrayBuffers, which are interpreted as containing UTF8 characters. When reading from HTTP responses, the response.body can be streamed directly to rewriter instances.
CSS Selectors
The html-rewriter library supports type, class, attribute, and ID CSS selectors, as well as child and descendent combinators.
You can, for example, rewrite a page to lazy load images, but load a hero image normally. In this example the hero image includes the ID “hero” so you can use :not pseudoclass with the ID selector img:not(#hero) to identify all of the non-hero images.
Here’s the JavaScript for your EdgeWorkers function.
If the index.html file contains the following details the hero image remains unchanged. Lazy load only applies to the other images.
The EdgeWorkers output will contain the following images.
Methods
onElement()
Registers a handler to run when a CSS selector matches. The handler takes an Element object as a parameter and provides functions to modify the document.
Handler Execution Ordering
When multiple handlers match an element, they run in the order that onHandler() was called.
The example below specifies the ordering.
The ordering occurs when the HTML runs.
The output in this example shows that the handler added on line one runs first, followed by the subsequent handlers.
Element Object
The Element object is an argument to the handler registered with the onElement()method. The handler calls functions on the Element to modify the output stream.
📘 You should not store the Element object. It’s reused when calling each handler. Using the object outside of the handler that it was passed into may have unexpected results.
Properties
selector
The CSS string passed to onElement().
tag
The lowercase name of the matched HTML tag.
Methods
The Element object supports methods for adding text around existing elements. The example below shows the output of processing <div>original</div>.
after()
Inserts new content immediately after the end tag of the matched element. The argument is the new text to insert.
When given the input <div></div>, the rewriter transforms it to <div></div>AFTER.
If the original document doesn’t include a close tag for the element, you can use the insert_implicit_close optional argument to differentiate between the element child and the appended text. For example, you can use the logic in the code sample below if you want to expand certain links, but </a> tags don’t appear reliably in the source document.
This lets you to gracefully handle malformed HTML with missing end tags.
The example below is rewritten so that the original <a> element now has an end tag.
append()
Inserts content at the end of the element.
This example adds scripts at the end of the <head> element.
<head></head> is re-written to the following.
before()
Inserts text immediately before the start tag of the matched element.
This example inserts a leading div before the title.
It transforms <h2 class='product-title'>Cheese slicer</h2> to a micro friendly format.
getAttribute()
Reads the value of an attribute name on the tag, returning undefined if the attribute does not exist.
This example changes the script path.
It uses the following input.
The rewriter changed the v1 in the path to v2.
prepend()
Inserts content right after the start tag of the element.
This example adds an onElement element to preload directives to a <head> element.
The rewriter changed <head></head> to the following.
removeAttribute()
Removes an attribute if it exists. Returns the value.
This example removes the background element of a body.
replaceChildren()
Removes the children of the current element and inserts content in place of them. Leaves the tags intact.
This example removes an inline script and instead loads a remote script.
Running on an input of <script>window.alert('hi')</script>, produces the following output.
replaceWith()
Removes the tags and element children. Inserts the passed content in its place.
This example registers a new callback.
It uses the following input.
The rewriter transforms the callback.
setAttribute()
Sets the value of the named attribute. Creates the attribute if one does not exist.
The following example.
When run on <div> produces the results below.
👍 A number of functions support an optional
TrailingOptargument. If the argument is present, the options object must include a property namedinsert_implicit_closewith a boolean value. When the value istrue, elements that are missing a close tag will have one inserted.
Insertion Ordering
When inserting text, the insertion point remains the same, even after other insertions. Consider a handler that uses multiple append statements.
The after() insertion point is the end of the close tag. This means that an input of <div></div> produces an output of <div></div>CBA.
- The
after()on line two inserts the A next to the close angle bracket. - The
after()on line three inserts the B between the close angle bracket and the previously inserted A. - The
after()on line four inserts the C between the close angle bracket and the previously inserted B.
Replacement and Nested Handlers
Handler matching is disabled during replacement.
For example, if you provide the following input.
With the following handlers.
The handler that matches on <i> will not run.
That means your handlers must not rely on side-effects such as, modifying variables in the module scope, when replacement is occurring.
Development Tips
When creating a new handler, it’s helpful to iterate on a controlled input. Rather than using an external document for your input, it’s possible to use a locally defined ReadableStream.
Here’s the expected output.
The ReadableStreams pattern is helpful when writing small integration tests.
📘 The
TextEncoderStreamis necessary because strings are written into the pipeline, rather than typed arrays.