feat: i18n: --lang translates report titles, headings, month names and the ui/web interfaces (#1025, #231) - #2736
Conversation
|
Please see https://github.com/acinader/hledger/blob/i18n-poc/doc/TRANSLATING.md as an entry point for reviewing this draft pr. |
0697094 to
2f50272
Compare
This comment was marked as off-topic.
This comment was marked as off-topic.
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring hledgerorg@8ddb03f changes into hledgerorg#2736 should bring hledgerorg#2736 upto parity with hledgerorg#2735 AI usage: drafted with Claude Code, reviewed and edited by the author. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring hledgerorg@8ddb03f changes into hledgerorg#2736 should bring hledgerorg#2736 upto parity with hledgerorg#2735 AI usage: drafted with Claude Code, reviewed and edited by the author. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring hledgerorg@8ddb03f changes into hledgerorg#2736 should bring hledgerorg#2736 upto parity with hledgerorg#2735 AI usage: drafted with Claude Code, reviewed and edited by the author.
b0251c0 to
5ef3a4e
Compare
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring hledgerorg@8ddb03f changes into hledgerorg#2736 should bring hledgerorg#2736 upto parity with hledgerorg#2735 AI usage: drafted with Claude Code, reviewed and edited by the author.
|
5ef3a4e is a good example to contemplate.
My current take is that we should merge in the i18n stuff, and maybe include the german po file as a starter? But I wouldn't claim that we have a german localization (l10n)? |
|
This is very comprehensive and impressive. Great work. We'll go with this for now. Some changes/discussion needed:
Bikeshedding:
Questions:
|
|
Using the english text as the keys is the the most conventional gettext approach I think. Which is generally a good thing. And it seems convenient to start with. But there will be churn when we want to tweak the ui messages. Is using a sum type of message ids, as in Henning's #2735, compatible with gettext tools ? |
|
Let's focus on just the library and CLI reports to begin with. |
|
Also let's defer the plural support, if not yet used (?)
Ah, but that's a good thing apparently. It signals staleness and updates needed. My claude session recommends combining a sum type for compiler checking plus english text for best translation experience:
|
|
Then again: just writing tr before string literals in our code is a very nice gentle first step. And we don't have a huge amount of text and haven't yet experienced the problems compiler checked ids would solve. I'm fine with the current approach for now. |
The journal and register handlers still set their titles with plain setTitle, so a French or German viewer got "journal - hledger-web" in the tab while the rest of the page was translated. The edit and upload pages already used setTitleI. Both titles are now catalog entries, with a translator note, and the German catalog has them. The browser test for German checks the tab titles too; since the English ones are lower case, the check cannot pass on an untranslated title. AI usage: Claude Fable 5.1, ~1k output tokens.
So that the cost of catalogs can be kept an eye on as they grow, --debug now reports, once, how long parsing the built-in catalogs took, and for each language loaded: where its catalog came from (built-in and/or the path of the user's file), how many entries it has, roughly how much text it holds, and how long loading and merging took. The size is measured after the timer stops, and none of it is computed without --debug. translations: parsed 1 built-in catalogs, 1.8 ms translations: loaded de (built-in, ~/.config/hledger/locale/de.po): 179 entries, ~8 KB of text, 0.5 ms The built-in parse is timed in availableLanguages, because listing the built-in languages is what forces it. AI usage: Claude Fable 5.1, ~2k output tokens.
quotedString collected every character of every quoted string into a String one at a time, then packed it; since nearly all of a catalog is quoted strings, that was most of the parse. It now takes runs of ordinary characters as slices of the input and handles only escapes a character at a time. A 920 KB catalog of 4000 entries parses in 36 ms instead of 137 ms (25 MB/s instead of 7), with a quarter of the transient allocation. The built-in German catalog goes from about 1.8 ms to about 0.9 ms. AI usage: Claude Fable 5.1, ~3k output tokens.
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring hledgerorg@8ddb03f changes into hledgerorg#2736 should bring hledgerorg#2736 upto parity with hledgerorg#2735 AI usage: Claude Fable 5.1, ~5k output tokens.
Add translations for the balance page AI usage: Claude Fable 5.1, ~6k output tokens.
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring hledgerorg@8ddb03f changes into hledgerorg#2736 should bring hledgerorg#2736 upto parity with hledgerorg#2735 AI usage: Claude Fable 5.1, ~5k output tokens.
|
Thanks for the updates. I'm going to merge this as-is if you don't object. I'm feeling the need to move things along, to possibly get preview 5 out on schedule (october 1st). |
|
I have a few revisions based on your comments that i am working on, so give
me a little more time, but I'm on board with your plan.
…On Tue, Sep 29, 2026 at 10:17 AM Simon Michael ***@***.***> wrote:
*simonmichael* left a comment (hledgerorg/hledger#2736)
<#2736 (comment)>
Thanks for the updates. I'm going to merge this as-is if you don't object.
I'm feeling the need to move things along, to possibly get preview 5 out on
schedule (october 1st).
—
Reply to this email directly, view it on GitHub
<#2736?email_source=notifications&email_token=AAFLBHC4VS3XCDIX3BFJXZ35RPVCJA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOBZGUYDSNJQGYYKM4TFMFZW63VGMF2XI2DPOKSWK5TFNZ2KYZTPN52GK4S7MNWGSY3L#issuecomment-5895095060>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/AAFLBHCTBYPLULXJAZGMSFD5RPVCJAVCNFSNUABDKJSXA33TNF2G64TZHM4TGMBRGQYTIO2JONZXKZJ3GU2DSMRXGA4TMNJRUF3AE>
.
Triage notifications, keep track of coding agent tasks and review pull
requests on the go with GitHub Mobile for iOS
<https://github.com/notifications/mobile/ios/AAFLBHDM32D573XJKBRC3TT5RPVCJA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOBZGUYDSNJQGYYKM4TFMFZW63VGMF2XI2DPOKSWK5TFNZ2KUZTPN52GK4S7NFXXG>
and Android
<https://github.com/notifications/mobile/android/AAFLBHGBYP4SIVZ32YOL3QL5RPVCJA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTKOBZGUYDSNJQGYYKM4TFMFZW63VGMF2XI2DPOKSWK5TFNZ2K4ZTPN52GK4S7MFXGI4TPNFSA>.
Download it today!
You are receiving this because you authored the thread.Message ID:
***@***.***>
|
|
Or would you rather get the web prs merged first ? |
The report renderers now bind tr, trc and trf to the report options once, in a where clause, and import Hledger.Utils.I18n qualified, so call sites read `tr "Balance Sheet"` rather than `tr translations_ "Balance Sheet"` (suggested in the hledgerorg#2736 review). The module docs now say when to mark a literal with i18n instead of translating it with tr: only for text stored in data, such as a report spec, before any catalog is at hand. AI usage: Claude Fable 5.1, ~2k output tokens.
|
AI usage. Done: every commit now carries the AI usage line with I have read I18n.hs; the note in the lib commit saying otherwise predated my My larger point is simply that gettext is a well understood interface that Plurals. I'd rather keep Naming. The CLI renderers now bind i18n versus tr. My overall conclusion is that this PR puts the machinery in place at low cost The implementation is also complete enough that if someone wanted to do a Finally, I'd like this to go in before my other PRs, as it will be easier on me |
|
Excellent. Thank you very much! |
"Monthly Balance Sheet" was an interval word plus a report name, joined by a space in the code. One adjective form had to fit every report, and the space could not be dropped. Now each compound report lists its own title per interval, so a translator renders the whole phrase as the language needs, with no grammar logic in Haskell. The German catalog grows from 14 to 44 title entries. English output is unchanged. this should bring 8ddb03f changes into #2736 should bring #2736 upto parity with #2735 AI usage: Claude Fable 5.1, ~5k output tokens.
|
I'm having trouble merging the web PRs, and I see why: the line numbers in the pot file make it very conflicty. This seems like it will be a problem. Could you change the tool to omit the line numbers - not quite as convenient for translators, but I think it's needed. |
|
Slightly related, I noticed that the texts from all --lang-supporting apps based on hledger-lib are now centralised in hledger-lib, ie hledger-lib has baked-in knowledge about things that depend on it, which is not good for modularity. But perhaps it is the pragmatic choice. |
The template's #: lines now name only the source files a text appears in. With line numbers, any change that shifted code rewrote hundreds of references on the next regeneration, so every branch that regenerated the template conflicted with every other (hledgerorg#2736 review). Now the template changes only when strings are added, removed, or move to another file. AI usage: Claude Opus 5.5, ~1k output tokens.
|
Well, that didn't go as smoothly as I was hoping! I've rebased and fiddled with the open web PRs so you should now be able to merge #2755 now writes the template's references as file names only, no line numbers, Catalogs in hledger-lib: You're right that it inverts the dependency. My |
|
Great, thanks.
Message ID: ***@***.***>
|
The template's #: lines now name only the source files a text appears in. With line numbers, any change that shifted code rewrote hundreds of references on the next regeneration, so every branch that regenerated the template conflicted with every other (#2736 review). Now the template changes only when strings are added, removed, or move to another file. AI usage: Claude Opus 5.5, ~1k output tokens.
|
It seems to me that --lang is not affecting hledger-web. |
|
You're browser is almost certainly sending I did not make --lang "win" by default and force all responses to be a particular locale. ?_LANG=xx -> cookie -> accept header -> --lang. http://127.0.0.1:5000/journal?_LANG=de switches to Derman and sets a cookie to remembers the choice. |







Translations for hledger's own output text, following the list discussion
(https://groups.google.com/g/hledger/c/PmcY3J8WiDc), #1025 and #231. German is
the second language.
What it does
hledger bs --lang de:--lang=LANG(usable in config files) selects atranslation catalog. It translates the compound reports' titles, section
titles and interval words, the balance and budget report titles and their
valuation descriptions, the Total/Average column headings and the Net: row
in text output, month names in period headings, and FODS sheet names.
?_LANG=de(remembered in a cookie), the
_LANGcookie, the browser's Accept-Language,then the server's
--lang. Pages, the add form and its validation messages,the help dialog and the file management pages are covered.
hledger-lib/locale/, embedded atbuild time. A user can override or add a language by dropping
~/.config/hledger/locale/LANG.poin place, without rebuilding: atranslator can work against a release binary.
--lang, every lookupis the identity. Checked by diffing the previous build's output for a few
dozen report invocations across output formats, and by the existing test
suites, which pass unchanged.
Design, against the points raised on the list
Stackage snapshot, and haskell-gettext last shipped in 2019. The mechanism
is one module,
Hledger.Utils.I18n, with no new dependency: a small POparser (including the Plural-Forms rule), lookup functions
tr,trc(contexts),
trf({placeholders}),trn(plurals), and language taghandling. The library is agent-generated (Claude Session) and
should be agent-maintained. Goal is to make implementation compliant
with existing tools for localization.
180 strings. The commits are split so that the hledger-lib and CLI parts
(commits 2 and 3) stand on their own; the hledger-ui and hledger-web
commits can be deferred.
--langplus a field on ReportOpts. The field,translations_, holds theloaded catalog rather than a tag, because the renderers are pure and the
catalog is loaded once in IO; the tag is inside it.
--lang=auto, which follows LANGUAGE, LC_ALL,LC_MESSAGES and LANG with gettext's rules, is implemented but opt-in; I am
happy to drop it from this PR if you would rather defer it.
the manual.
--titleand--subreport-titleswin over translation, and number formatsare untouched (they come from the journal).
--titlethey are translated in csv/tsv/json title rows too; column headings in
those formats stay English. If you would rather csv stay entirely English,
the alternative is keeping titles structured until rendering, as
translate report texts to user language using haskell-gettext #2735
does; that is a bigger refactor and can follow.
One deliberate divergence: PO catalogs rather than Shakespeare's
.msgfiles.hledger-web does use Shakespeare's I18N interface (
RenderMessage,_{...}in templates), with these catalogs behind it. PO was chosen because complete
translations come from translators, and PO is what their tools (Poedit,
Weblate) speak: it gives contexts, translator notes, a fuzzy workflow, and
lets someone test a catalog against a release without a Haskell build. My POV being
that gettext is mature and we can reliably maintain the functionality we want
with agents maintaing the i18n core code.
Docs and tooling
options list; hledger-web's manual: a Language section.
doc/TRANSLATING.md: a step-by-step guide for translators who are neitherprogrammers nor professionals, with a worked example.
tools/i18n-extract.pyandjust i18n-pot,i18n-check,i18n-merge,i18n-pseudomaintain the catalogs.Tests
hledger functional suite (9 new cases in
hledger/test/i18n.test),hledger-lib unit tests for the parser, plural rules and tag handling,
hledger-web yesod tests for language selection, the cookie and a hostile
catalog rendered as text, and a Playwright spec for the German UI. All four
packages build warning-free.
Caveats
I have not personally walked through the steps indoc/TRANSLATING.mdasa translator would; the guide was checked mechanically (a throwaway French
catalog placed in the config directory, exercised through
--lang fr), notby a person following it cold.
No performance testing. With no
--langoption nothing is parsed or read,so the default path should be unaffected. Any
--langvalue, includingenandauto, lists the override directory and parses the built-incatalogs (a few hundred entries). I have not measured either path.
Follow-ups
hledger_site symlink and SUMMARY entry, which I will open separately.
code contributions that touch localizable strings.
Not in this PR
The hledger-ui help dialog and register labels (need a layout rework),
statslabel widths, html/fods column headings, and translating errormessages or manuals.
Terminology
The German catalog uses the terms Henning chose in #2735 (Einnahmen,
Ausgaben, Einnahmenüberschussrechnung, Vermögen, Gesamt, Überschuss,
Einheit), so the two agree; see
hledger-lib/locale/README.md. The oneplace I kept a different word is hledger-web's account column (Konto).
examples/i18n/de.journalnames its accounts aktiva and passiva, whichGerman speakers may want to reconcile.
AI usage: designed and implemented with Claude Code, reviewed, tested, and edited by the author.
🤖 Generated with Claude Code
https://claude.ai/code/session_01G2VXnprjHmXZR8tgWPV3vz