Automattic\WooCommerce\EmailEditor\Integrations\Utils

Html_Processing_Helper::sanitize_caption_html │ public static │ WC 1.0

Sanitize caption HTML to allow only specific tags and attributes.

Method of the class: Html_Processing_Helper{}

No Hooks.

Returns

string. Sanitized caption HTML.

Usage

$result = Html_Processing_Helper::sanitize_caption_html( $caption_html ): string;
$caption_html(string) (required)
Raw caption HTML.

Html_Processing_Helper::sanitize_caption_html() code WC 11.1.2

public static function sanitize_caption_html( string $caption_html ): string {
	// If no HTML tags, return as-is.
	if ( false === strpos( $caption_html, '<' ) ) {
		return $caption_html;
	}

	/*
	 * Remove executable elements together with their content. The allow-list
	 * pass below keeps the text of the tags it strips, which would turn a
	 * script body into visible caption text.
	 *
	 * One pattern per element rather than a single alternation that backreferences
	 * the opening name: with the closing tag spelled out as a literal, PCRE can
	 * reject a caption that never closes the element outright, instead of scanning
	 * to the end of the caption once for every opening tag it finds.
	 */
	foreach ( array( 'script', 'style', 'iframe', 'object', 'embed', 'form', 'input', 'button' ) as $tag ) {
		$result = preg_replace( '/<' . $tag . '\b[^>]*>.*?<\/' . $tag . '>/is', '', $caption_html );
		// A caption long enough to exhaust PCRE's limits keeps whatever the pass
		// managed to remove. wp_kses() below still drops the element itself, so
		// only the text that was inside it is left behind.
		if ( null !== $result ) {
			$caption_html = $result;
		}
	}

	/*
	 * Reduce the markup to the allowed elements and attributes.
	 *
	 * wp_kses() rebuilds each tag it keeps from the allow-list rather than
	 * passing the authored text through, which is what makes malformed markup
	 * safe here. A ">" inside an attribute value still ends the tag span early,
	 * but the allowed tag that falls out of that split is re-emitted with only
	 * its allowed attributes. A tag left without a closing ">" cannot survive as
	 * markup either, though which way it goes depends on what is attached to the
	 * pre_kses hook: core escapes it to text there, and without that callback
	 * kses drops the tag or supplies the missing ">" itself.
	 */
	$caption_html = wp_kses( $caption_html, self::get_allowed_caption_html(), array( 'http', 'https', 'mailto', 'tel' ) );

	/*
	 * Narrow the attribute values that survived: protocol allow-list on href,
	 * safe CSS properties on style, rel and target normalization. Every
	 * remaining tag is allowed, so every attribute on it is validated.
	 */
	$html = new \WP_HTML_Tag_Processor( $caption_html );
	while ( $html->next_tag() ) {
		$attributes = $html->get_attribute_names_with_prefix( '' );
		if ( is_array( $attributes ) ) {
			foreach ( $attributes as $attr_name ) {
				self::validate_caption_attribute( $html, $attr_name );
			}
		}
	}

	return $html->get_updated_html();
}