Automattic\WooCommerce\EmailEditor\Integrations\Utils
Html_Processing_Helper::sanitize_caption_html
Sanitize caption HTML to allow only specific tags and attributes.
Method of the class: Html_Processing_Helper{}
No Hooks.
Returns
string. Sanitized caption HTML.
Usage
$result = Html_Processing_Helper::sanitize_caption_html( $caption_html ): string;
- $caption_html(string) (required)
- Raw caption HTML.
Html_Processing_Helper::sanitize_caption_html() Html Processing Helper::sanitize caption html code WC 11.1.2
public static function sanitize_caption_html( string $caption_html ): string {
// If no HTML tags, return as-is.
if ( false === strpos( $caption_html, '<' ) ) {
return $caption_html;
}
/*
* Remove executable elements together with their content. The allow-list
* pass below keeps the text of the tags it strips, which would turn a
* script body into visible caption text.
*
* One pattern per element rather than a single alternation that backreferences
* the opening name: with the closing tag spelled out as a literal, PCRE can
* reject a caption that never closes the element outright, instead of scanning
* to the end of the caption once for every opening tag it finds.
*/
foreach ( array( 'script', 'style', 'iframe', 'object', 'embed', 'form', 'input', 'button' ) as $tag ) {
$result = preg_replace( '/<' . $tag . '\b[^>]*>.*?<\/' . $tag . '>/is', '', $caption_html );
// A caption long enough to exhaust PCRE's limits keeps whatever the pass
// managed to remove. wp_kses() below still drops the element itself, so
// only the text that was inside it is left behind.
if ( null !== $result ) {
$caption_html = $result;
}
}
/*
* Reduce the markup to the allowed elements and attributes.
*
* wp_kses() rebuilds each tag it keeps from the allow-list rather than
* passing the authored text through, which is what makes malformed markup
* safe here. A ">" inside an attribute value still ends the tag span early,
* but the allowed tag that falls out of that split is re-emitted with only
* its allowed attributes. A tag left without a closing ">" cannot survive as
* markup either, though which way it goes depends on what is attached to the
* pre_kses hook: core escapes it to text there, and without that callback
* kses drops the tag or supplies the missing ">" itself.
*/
$caption_html = wp_kses( $caption_html, self::get_allowed_caption_html(), array( 'http', 'https', 'mailto', 'tel' ) );
/*
* Narrow the attribute values that survived: protocol allow-list on href,
* safe CSS properties on style, rel and target normalization. Every
* remaining tag is allowed, so every attribute on it is validated.
*/
$html = new \WP_HTML_Tag_Processor( $caption_html );
while ( $html->next_tag() ) {
$attributes = $html->get_attribute_names_with_prefix( '' );
if ( is_array( $attributes ) ) {
foreach ( $attributes as $attr_name ) {
self::validate_caption_attribute( $html, $attr_name );
}
}
}
return $html->get_updated_html();
}