Comment out the header line - #12
Merged
Merged
Conversation
Parsers that treat the file as pure data read `Abbreviation<TAB>SMILES` as a row and try to parse "SMILES" as a structure. Prefixing the header with `#` makes it skippable by the usual convention. index.html strips a leading `#` from the header line so the viewer still renders two named columns. lint.py's pandas round-trip and the duplicate check are unaffected.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
abbreviations.smiline 1 becomes# Abbreviation<TAB>SMILES.Consumers that treat the file as pure data currently read the header as a row and try to parse the literal string
SMILESas a structure. In BChemXtract that surfaces as aWARN Invalid SMILESwith aCDKExceptionstack trace on every JVM start.#is the comment convention those parsers already skip.Why the second file
index.htmluses line 1 for the DataTables column titles and splits it on whitespace, so a bare#would have produced three columns and misaligned the data. It now strips a leading#from the header line before splitting — the viewer rendersAbbreviation/SMILESexactly as before.Verification
pd.read_csv(sep='\t')→ columns['# Abbreviation', 'SMILES'],to_csv(sep='\t', index=False)round-trip byte-identical, solint.pypasses./scripts/check_duplicates_abbreviations.shpasses;cut -f1 | sort | uniq -dalso cleanindex.htmlparse re-run in node: headers["Abbreviation","SMILES"], 300 data rows, last rowNHMe → [*]NCNo data rows were added, removed, or changed.
🤖 Generated with Claude Code