Skip to content

Integrate fragment parser options with trusted types - #12583

Open
noamr wants to merge 4 commits into
mainfrom
noamr/cpo
Open

noamr wants to merge 4 commits into
mainfrom
noamr/cpo

Conversation

@noamr

@noamr noamr commented Jun 16, 2026 •

Copy link
Copy Markdown
Contributor

In the places where Trusted Types' createHTML is called, we now also check for createParserOptions and call that if it's available, by calling "get trusted type compliant input" instead of "get trusted type compliant string".

This now allows a default policy to supply a sanitizer or change script-running behavior for the legacy markup insertion methods (innerHTML, outerHTML, insertAdjacentHTML(), createContextualFragment()), and to alter the sanitizer/script behavior of modern calls like setHTMLUnsafe() and parseHTMLUnsafe().

The exceptions to this are srcdoc, document.write(), DOMParser.parseFromString(), and the insertHTML execCommand. We can resolve separately whether these should also be sanitized with createParserOptions.

Structural changes

  • The safe boolean and scriptingMode parameters of "set and filter HTML" / "fragment parsing algorithm steps" are replaced by a single fragment parser mode enum: Safe (setHTML(), parseHTML()), Unsafe (setHTMLUnsafe(), parseHTMLUnsafe()), and Legacy (innerHTML, outerHTML, insertAdjacentHTML(), createContextualFragment()). Safe and Unsafe always use the HTML parser and allow declarative shadow roots; Legacy keeps today's behavior (XML parser in XML documents, no declarative shadow roots).

  • Since all markup insertion methods can now include a sanitizer, most of the steps from "set and filter HTML" are folded into the "fragment parsing algorithm steps" (which moved next to innerHTML), with handling of a null sanitizer when appropriate.

  • "get trusted type compliant input" takes an optionsFromAuthor boolean. It is true for author-supplied options (setHTMLUnsafe(), parseHTMLUnsafe()) and false when the sink constructs the options itself (the legacy sinks, including createContextualFragment(), which sets runScripts to true). This lets Trusted Types throw when TrustedHTML is combined with unvetted author options that need vetting (runScripts: true or a non-empty sanitizer), while not throwing for sink-constructed options.

Observable changes beyond the Trusted Types integration

  • The = {} defaults are removed from SetHTMLUnsafeOptions.sanitizer and ParseHTMLUnsafeOptions.sanitizer so that a default policy can distinguish an omitted sanitizer member from an explicit configuration.

  • A Sanitizer object passed to setHTML()/parseHTML() is no longer mutated: the sink now works on a fresh copy of its configuration, so sanitizer.get() is unchanged after use. Previously the "remove unsafe" step mutated the author's object in place.

  • "configure a sanitizer" clones its input configuration before canonicalizing, fixing an issue where the built-in safe default configuration singleton was mutated in place on first use.

  • createContextualFragment() now creates its fallback body context element in the range's start node's node document (previously "this's node document", but a Range has no node document).

  • A default policy is only allowed to provide a sanitizer for HTML documents. In XML documents the createParserOptions shortcut does not apply, so createHTML (or a TrustedHTML) is still required, exactly as today. (setHTMLUnsafe() in an XHTML document still uses the HTML parser and therefore does honor createParserOptions.)

Together with w3c/trusted-types#606, which must land first (this PR links to anchors defined there); more details on the Trusted Types side are described there.

(See WHATWG Working Mode: Changes for more details.)


/dom.html ( diff )
/dynamic-markup-insertion.html ( diff )
/iframe-embed-object.html ( diff )
/infrastructure.html ( diff )
/parsing.html ( diff )
/scripting.html ( diff )

@noamr noamr changed the title Integrate createParserOptions (draft) Integrate createParserOptions Jul 13, 2026
@noamr noamr changed the title Integrate createParserOptions Integrate fragment parser options with trusted types Jul 13, 2026
@noamr noamr closed this Jul 13, 2026
@noamr noamr reopened this Jul 13, 2026
@noamr
noamr marked this pull request as ready for review July 13, 2026 19:43
@noamr
noamr requested review from annevk, lukewarlow and zcorpan July 13, 2026 19:46
@noamr noamr added the agenda+ To be discussed at a triage meeting label Jul 15, 2026
@lukewarlow

Copy link
Copy Markdown
Member

Two things for the top comment, I don't think you do fold set and filter html anymore?

Also this also doesn't touch execCommand with the insertHTML command. (This is probably the right move because that isnt specced correctly to start with. But worth calling out probably).

Comment thread source Outdated
@noamr

noamr commented Jul 17, 2026

Copy link
Copy Markdown
Contributor Author

Two things for the top comment, I don't think you do fold set and filter html anymore?

Also this also doesn't touch execCommand with the insertHTML command. (This is probably the right move because that isnt specced correctly to start with. But worth calling out probably).

Thanks, OP updated.

@lukewarlow

Copy link
Copy Markdown
Member

Did we decide it was okay to not enforce the sanitizing if the document was XML?

@noamr

noamr commented Jul 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Did we decide it was okay to not enforce the sanitizing if the document was XML?

That's what I understood but @mozfreddyb, @evilpie or @otherdaniel would know more. Sanitization is not specified for the XML parser.
Note that this is the status quo that is not changed here - what this does is allows you to sanitize when setting HTML with one of the old methods (and soon with streaming).

@otherdaniel

Copy link
Copy Markdown
Contributor

Did we decide it was okay to not enforce the sanitizing if the document was XML?

That's what I understood but @mozfreddyb, @evilpie or @otherdaniel would know more. Sanitization is not specified for the XML parser. Note that this is the status quo that is not changed here - what this does is allows you to sanitize when setting HTML with one of the old methods (and soon with streaming).

What I remember is that we specified Sanitizer API only for methods that would inherently only support HTML syntax. setHTML, setHTMLUnsafe, parseHTML, and parseHTMLUnsafe were all new methods, that simply don't support XML syntax. It's right there in the names. We didn't say what to do for XML syntax because we didn't have to.

This is probably best articulated in the hopelessly outdated explainer

Nearly all interesting bits are specified in terms of DOM & DOM operations, so I'd expect this to be easy to adapt to XML. But IMHO, application to XML-parser parse trees requires a second look, since it invalidates one of the assumptions we had when specifying any of this.


A silly example, but the only one I can think of: CDataSection in https://wicg.github.io/sanitizer-api/#sanitize-core step 1.1. That shouldn't be difficult to fix; but at least for now Sanitizer would assert-fail on (some) XML parse trees. In our implementation, there's a runtime assert there.

@lukewarlow

Copy link
Copy Markdown
Member

I guess my main concern is people defining a trusted types policy with this new function thinking it protects them and then it doesn't because they're in XHTML or something?

Assuming I'm reading this right you'd end up with a default policy explicitly setup to remove unsafe and then it actually no-ops when it's called by a legacy sync in XML.

@noamr

noamr commented Jul 17, 2026

Copy link
Copy Markdown
Contributor Author

I guess my main concern is people defining a trusted types policy with this new function thinking it protects them and then it doesn't because they're in XHTML or something?
Assuming I'm reading this right you'd end up with a default policy explicitly setup to remove unsafe and then it actually no-ops when it's called by a legacy sync in XML.

You mean sink?

Yea it's limited in that way. But createHTML is still there... until we have some solution for this people should probably still use both in their policy or protect XML in other means.

@noamr

noamr commented Jul 17, 2026

Copy link
Copy Markdown
Contributor Author

I guess my main concern is people defining a trusted types policy with this new function thinking it protects them and then it doesn't because they're in XHTML or something?
Assuming I'm reading this right you'd end up with a default policy explicitly setup to remove unsafe and then it actually no-ops when it's called by a legacy sync in XML.

You mean sink?

Yea it's limited in that way. But createHTML is still there... until we have some solution for this people should probably still use both in their policy or protect XML in other means.

I think that the specific guidance to developers to be to check the type of document when they create the default policy, use createHTML with the appropriate userland sanitizer if either this is an XML document or TrustedParserOptions is not supported, and createParserOptions otherwise

@lukewarlow

Copy link
Copy Markdown
Member

Non-authoratative LGTM. I'm still slightly unsure about the XML case mentioned above but if the consensus is that it's fine then I buy that.

@noamr noamr closed this Jul 23, 2026
@noamr noamr reopened this Jul 23, 2026
@noamr noamr removed the agenda+ To be discussed at a triage meeting label Jul 23, 2026
@noamr noamr closed this Aug 4, 2026
@noamr noamr reopened this Aug 4, 2026
@noamr

noamr commented Aug 28, 2026

Copy link
Copy Markdown
Contributor Author

It seems that createHTML in a (default) policy is mandatory whereas createParserOptions is optional. I think that make some sense for existing code, but it also means that there's no way to create a TT default policy that only has createParserOptions, which I think could be a bug. A workaround would be to use a policy that has createHTML as the identity function, but then they'd open up srcdoc/document.write/etc, which is a clear security bug.
The reason I am pressing on this is that I think Trusted Type has a semantic issue where createHTML can't easily perform context-aware sanitization (because it mostly sees the HTML string, not the HTML tree or the parsing state). A more robust createParserOptions can fix this issue with Trusted Types. Ideally, we'd want people to use TT with a built-in browser sanitizer than user-space HTML parsing/sanitization code.
Could we invert the lookup and make policies createParserOptions work without requiring createHTML()?

I think that's fine. To make sure I understand - the idea is that a required HTML TT in your CSP would reach out to createParserOptions, and if not reach out to createHTML, but fail if only if both are missing?

I think I'm OK with that. @otherdaniel

Updated in w3c/trusted-types#606

@noamr noamr closed this Aug 28, 2026
@noamr noamr reopened this Aug 28, 2026

@zcorpan zcorpan left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Opus+Codex review of new changes

Comment thread source Outdated
Comment thread source Outdated
Comment thread source Outdated
Comment thread source
@evilpie

evilpie commented Sep 1, 2026

Copy link
Copy Markdown
Member

LGTM. The way the options are handled across TT is now quite complicated, but from what I can tell the behavior is still correct.

@noamr

noamr commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

Applied multiple comments from offline review by @annevk.
Most notably, making a report-only csp check when it's an XML doc, exporting a few dfns for trusted types, updated the domintro to mention some missing error conditions. The corresponding trusted types PR is updated as well to handle cloning in a more explicit way, and to carry over the input sanitizer when a TrustedHTML object is passed as input.

@noamr
noamr force-pushed the noamr/cpo branch 2 times, most recently from 810d7a7 to 02b7179 Compare September 3, 2026 20:52

@annevk annevk left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, but might need some changes given my latest comments here: w3c/trusted-types#606 (review).

Comment thread source Outdated
noamr and others added 2 commits September 29, 2026 11:53
createParserOptions

TrustedParseOptions in IDL

Still wrap things in set and filter HTML

Add clarification about XMLL vs HTML

Some fixes from ChatGPT

Throw when applying sanitizer to XML document in a legacy method

Rebaseline to new TT changes

Handle default sanitizer correctly when safe is false, and fix TT throwIfMissing polarity

Revert "Handle default sanitizer correctly when safe is false, and fix TT throwIfMissing polarity"

This reverts commit 7898aca.

Handle null in TrustedParserOptions

Fix ccf

Add some null checks

Address review comments for PR 12583 (TrustedParserOptions & sanitizer options)

Address Opus review notes

Improve scripting mode switchh

nit

Fix unclosed li tags in script preparation and fragment parsing

Use <span> in IDL

Fix <span> in WebIDL

Remove <code> in WebIDL

Fix WebIDL type xref and line wrapping in parseHTMLUnsafe and fragment parsing

Editorial: fix phrasing, line wrapping, and formatting in sanitizer options

Editorial: fix phrasing, line wrapping, and formatting in sanitizer options

nits

Address review comments on TrustedParserOptions & sanitizer options

specfmt

Export 'canonicalize the configuration' and 'built-in safe default configuration'

Editorial: Remove unused ParseHTMLUnsafeOptions from fragment parsing algorithm steps signature

Editorial: address review comments on CPO PR

Editorial: address style and markup reviews on CPO

Editorial: address review comments on CPO PR (B4, B9, N5, N6, N8)

Address review comments on parser options and trusted types

More explicit naming

Pass node document's type to get trusted type compliant input

Editorial: address review comments on TT and XML handling

Remove comment

Editorial: Remove inaccurate domintro sentences about XML TypeError

Editorial: note that sanitization is not supported for XML documents

TrustedParserOptions->TrustedHTMLParserOptions

Editorial: Rename sanitizerSpec to sanitizerInput in get a sanitizer instance from options

Address review feedback on fragment parser options and trusted types
This lets Trusted Types reject author-supplied `runScripts` without breaking `createContextualFragment`. See w3c/trusted-types#616.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@noamr

noamr commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

I think this and w3c/trusted-types#606 are ready to land.

@zcorpan

zcorpan commented Oct 2, 2026

Copy link
Copy Markdown
Member
AI review

Line numbers are source at 8cfede2.

  • source:126248-126252, 126265-126269, 126292-126296 (domintro): the TypeError sentence is wrong. setHTMLUnsafe("<b>", {}) under enforcement with a default policy that only has createHTML doesn't throw. The options don't need vetting, so it falls through to createHTML. It only throws when neither path vets the input, or when author options need vetting and nothing vets them. It also misses the cases where createParserOptions throws or returns null.
  • source:126566, 126571 (wording): the mode definitions phrase the declarative shadow roots flag inconsistently ("while … flag is true" vs "and declarative shadow roots are allowed (the … flag is true)"). This also carries into html#12758.
  • The script note added near source:67740: it says scripts inserted via innerHTML/outerHTML/insertAdjacentHTML run if a default policy enables runScripts. The text is accurate, but this behavior plus the fact that createParserOptions bypasses createHTML (index.bs:1461-1462) would be worth a line in the Security Considerations. A default policy that does createParserOptions: o => o alongside a DOMPurify-based createHTML silently stops sanitizing innerHTML.

The this's → node's node document fix in createContextualFragment is good: Range has no node document.

@noamr

noamr commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

AI review
Line numbers are source at 8cfede2.

  • source:126248-126252, 126265-126269, 126292-126296 (domintro): the TypeError sentence is wrong. setHTMLUnsafe("<b>", {}) under enforcement with a default policy that only has createHTML doesn't throw. The options don't need vetting, so it falls through to createHTML. It only throws when neither path vets the input, or when author options need vetting and nothing vets them. It also misses the cases where createParserOptions throws or returns null.
  • source:126566, 126571 (wording): the mode definitions phrase the declarative shadow roots flag inconsistently ("while … flag is true" vs "and declarative shadow roots are allowed (the … flag is true)"). This also carries into html#12758.
  • The script note added near source:67740: it says scripts inserted via innerHTML/outerHTML/insertAdjacentHTML run if a default policy enables runScripts. The text is accurate, but this behavior plus the fact that createParserOptions bypasses createHTML (index.bs:1461-1462) would be worth a line in the Security Considerations. A default policy that does createParserOptions: o => o alongside a DOMPurify-based createHTML silently stops sanitizing innerHTML.

The this's → node's node document fix in createContextualFragment is good: Range has no node document.

Fixed all of the above

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

do not merge yet Pull request must not be merged per rationale in comment

Development

Successfully merging this pull request may close these issues.

7 participants