Skip to content

fix(renderer): skip a body tag inside head scripts when measuring body text - #606

Merged
us merged 1 commit into
mainfrom
fix/body-text-len-body-tag-in-script
Oct 4, 2026
Merged

us merged 1 commit into
mainfrom
fix/body-text-len-body-tag-in-script

Conversation

@us

@us us commented Oct 4, 2026

Copy link
Copy Markdown
Member

ref #605

html_body_text_len found the opening tag with find("<body"), which matches the first mention anywhere in the document, including inside a <script> in <head>. A loader that writes <body><script src=x><\/script></body> into an iframe made the real body measure 0 characters, so a fully loaded page was classified as a thin render and escalated to the next tier.

  • New find_body_open skips comments and the content of script, style, template and textarea, and steps over whole tags so a <script or <!-- inside a quoted attribute value is not read as a raw-text region.
  • Falls back to the previous plain find("<body") when the scan finds nothing (unterminated comment or raw-text element).
  • The closing-tag search and the counting loop are unchanged.
  • Tests: the reproduction from the issue (0 before, 17 after), raw-text and comment skipping, case-insensitive tags, <bodyguard>, attribute-value lookalikes, empty comments, unterminated cases.

Checked against the page from the issue, rendered HTML of about 1.16 MB: the first <body sits inside the head script (the measured span was 49 characters, all tags), the scan now lands on the real body.

No config, API or response shape changes. Pages that were measured correctly before measure the same; only pages whose first <body mention is not the real tag change.

…y text

html_body_text_len located the opening tag with a plain find("<body"), which
matches the first mention anywhere in the document. A script in <head> that
writes iframe markup (`<body><script src=x><\/script></body>`) made the real
body measure 0 characters, so a fully loaded page was classified as a thin
render and escalated to the next tier.

- add find_body_open, which skips comments and the content of script, style,
  template and textarea elements, and steps over whole tags so a `<script` in
  a quoted attribute value is not read as a raw-text region
- fall back to the previous plain find when the scan finds nothing
- the closing-tag search and the counting loop are unchanged

ref #605
@us
us merged commit baacc11 into main Oct 4, 2026
10 checks passed
@github-actions github-actions Bot locked and limited conversation to collaborators Oct 4, 2026
@us
us deleted the fix/body-text-len-body-tag-in-script branch October 4, 2026 16:13
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant