fix(renderer): skip a body tag inside head scripts when measuring body text - #606
Merged
Merged
Conversation
…y text
html_body_text_len located the opening tag with a plain find("<body"), which
matches the first mention anywhere in the document. A script in <head> that
writes iframe markup (`<body><script src=x><\/script></body>`) made the real
body measure 0 characters, so a fully loaded page was classified as a thin
render and escalated to the next tier.
- add find_body_open, which skips comments and the content of script, style,
template and textarea elements, and steps over whole tags so a `<script` in
a quoted attribute value is not read as a raw-text region
- fall back to the previous plain find when the scan finds nothing
- the closing-tag search and the counting loop are unchanged
ref #605
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ref #605
html_body_text_lenfound the opening tag withfind("<body"), which matches the first mention anywhere in the document, including inside a<script>in<head>. A loader that writes<body><script src=x><\/script></body>into an iframe made the real body measure 0 characters, so a fully loaded page was classified as a thin render and escalated to the next tier.find_body_openskips comments and the content ofscript,style,templateandtextarea, and steps over whole tags so a<scriptor<!--inside a quoted attribute value is not read as a raw-text region.find("<body")when the scan finds nothing (unterminated comment or raw-text element).<bodyguard>, attribute-value lookalikes, empty comments, unterminated cases.Checked against the page from the issue, rendered HTML of about 1.16 MB: the first
<bodysits inside the head script (the measured span was 49 characters, all tags), the scan now lands on the real body.No config, API or response shape changes. Pages that were measured correctly before measure the same; only pages whose first
<bodymention is not the real tag change.