Skip to content

feat: i18n: --lang translates report titles, headings, month names and the ui/web interfaces (#1025, #231) - #2736

Merged
simonmichael merged 13 commits into
hledgerorg:mainfrom
acinader:i18n-poc
Sep 29, 2026
Merged

simonmichael merged 13 commits into
hledgerorg:mainfrom
acinader:i18n-poc

Conversation

@acinader

@acinader acinader commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Translations for hledger's own output text, following the list discussion
(https://groups.google.com/g/hledger/c/PmcY3J8WiDc), #1025 and #231. German is
the second language.

What it does

hledger bs --lang de:

Bilanz 2008-12-31

                    || 2008-12-31
====================++============
 Vermögen           ||
--------------------++------------
 assets:bank:saving ||         $1
 assets:cash        ||        $-2
--------------------++------------
                    ||        $-1
====================++============
 Verbindlichkeiten  ||
--------------------++------------
 liabilities:debts  ||        $-1
--------------------++------------
                    ||        $-1
====================++============
 Überschuss:        ||          0
  • A new general option --lang=LANG (usable in config files) selects a
    translation catalog. It translates the compound reports' titles, section
    titles and interval words, the balance and budget report titles and their
    valuation descriptions, the Total/Average column headings and the Net: row
    in text output, month names in period headings, and FODS sheet names.
  • hledger-ui translates its menu and screen names.
  • hledger-web serves each request in the viewer's language: ?_LANG=de
    (remembered in a cookie), the _LANG cookie, the browser's Accept-Language,
    then the server's --lang. Pages, the add form and its validation messages,
    the help dialog and the file management pages are covered.
  • Translations are gettext PO files in hledger-lib/locale/, embedded at
    build time. A user can override or add a language by dropping
    ~/.config/hledger/locale/LANG.po in place, without rebuilding: a
    translator can work against a release binary.
  • English output is unchanged byte for byte: with no --lang, every lookup
    is the identity. Checked by diffing the previous build's output for a few
    dozen report invocations across output formats, and by the existing test
    suites, which pass unchanged.

Design, against the points raised on the list

  • No i18n library. haskell-gettext, hgettext and i18n are all absent from the
    Stackage snapshot, and haskell-gettext last shipped in 2019. The mechanism
    is one module, Hledger.Utils.I18n, with no new dependency: a small PO
    parser (including the Plural-Forms rule), lookup functions tr, trc
    (contexts), trf ({placeholders}), trn (plurals), and language tag
    handling. The library is agent-generated (Claude Session) and
    should be agent-maintained. Goal is to make implementation compliant
    with existing tools for localization.
  • Reports first, complete translations. The German catalog translates all
    180 strings. The commits are split so that the hledger-lib and CLI parts
    (commits 2 and 3) stand on their own; the hledger-ui and hledger-web
    commits can be deferred.
  • --lang plus a field on ReportOpts. The field, translations_, holds the
    loaded catalog rather than a tag, because the renderers are pure and the
    catalog is loaded once in IO; the tag is inside it.
  • English by default. --lang=auto, which follows LANGUAGE, LC_ALL,
    LC_MESSAGES and LANG with gettext's rules, is implemented but opt-in; I am
    happy to drop it from this PR if you would rather defer it.
  • Config files: documented, with an example, in the new Languages section of
    the manual.
  • --title and --subreport-titles win over translation, and number formats
    are untouched (they come from the journal).
  • Titles are report data shared by every output format, so like --title
    they are translated in csv/tsv/json title rows too; column headings in
    those formats stay English. If you would rather csv stay entirely English,
    the alternative is keeping titles structured until rendering, as
    translate report texts to user language using haskell-gettext #2735
    does; that is a bigger refactor and can follow.

One deliberate divergence: PO catalogs rather than Shakespeare's .msg files.
hledger-web does use Shakespeare's I18N interface (RenderMessage, _{...}
in templates), with these catalogs behind it. PO was chosen because complete
translations come from translators, and PO is what their tools (Poedit,
Weblate) speak: it gives contexts, translator notes, a fuzzy workflow, and
lets someone test a catalog against a release without a Haskell build. My POV being
that gettext is mature and we can reliably maintain the functionality we want
with agents maintaing the i18n core code.

Docs and tooling

  • Manual: a Languages section, the environment variables, the general
    options list; hledger-web's manual: a Language section.
  • doc/TRANSLATING.md: a step-by-step guide for translators who are neither
    programmers nor professionals, with a worked example.
  • tools/i18n-extract.py and just i18n-pot, i18n-check, i18n-merge,
    i18n-pseudo maintain the catalogs.

Tests

hledger functional suite (9 new cases in hledger/test/i18n.test),
hledger-lib unit tests for the parser, plural rules and tag handling,
hledger-web yesod tests for language selection, the cookie and a hostile
catalog rendered as text, and a Playwright spec for the German UI. All four
packages build warning-free.

Caveats

  • I have not personally walked through the steps in doc/TRANSLATING.md as
    a translator would; the guide was checked mechanically (a throwaway French
    catalog placed in the config directory, exercised through --lang fr), not
    by a person following it cold.
  • No performance testing. With no --lang option nothing is parsed or read,
    so the default path should be unaffected. Any --lang value, including
    en and auto, lists the override directory and parses the built-in
    catalogs (a few hundred entries). I have not measured either path.

Follow-ups

  • The manual links to hledger.org/TRANSLATING.html; that needs the usual
    hledger_site symlink and SUMMARY entry, which I will open separately.
  • Skills and documentation to help Agents assist with translation and
    code contributions that touch localizable strings.

Not in this PR

The hledger-ui help dialog and register labels (need a layout rework),
stats label widths, html/fods column headings, and translating error
messages or manuals.

Terminology

The German catalog uses the terms Henning chose in #2735 (Einnahmen,
Ausgaben, Einnahmenüberschussrechnung, Vermögen, Gesamt, Überschuss,
Einheit), so the two agree; see hledger-lib/locale/README.md. The one
place I kept a different word is hledger-web's account column (Konto).
examples/i18n/de.journal names its accounts aktiva and passiva, which
German speakers may want to reconcile.

AI usage: designed and implemented with Claude Code, reviewed, tested, and edited by the author.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G2VXnprjHmXZR8tgWPV3vz

@acinader

acinader commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Please see https://github.com/acinader/hledger/blob/i18n-poc/doc/TRANSLATING.md as an entry point for reviewing this draft pr.

@acinader
acinader force-pushed the i18n-poc branch 3 times, most recently from 0697094 to 2f50272 Compare September 17, 2026 23:27
@acinader

acinader commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

I don't speak French, so it's hard to tell....

Poedit tool for editing translation files.

image

Some screen shots

image image image image image image

pretty cool.

@acinader

This comment was marked as off-topic.

acinader added a commit to acinader/hledger that referenced this pull request Sep 20, 2026
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring hledgerorg@8ddb03f changes into hledgerorg#2736

should bring hledgerorg#2736 upto parity with hledgerorg#2735

AI usage: drafted with Claude Code, reviewed and edited by the author.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
acinader added a commit to acinader/hledger that referenced this pull request Sep 20, 2026
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring hledgerorg@8ddb03f changes into hledgerorg#2736

should bring hledgerorg#2736 upto parity with hledgerorg#2735

AI usage: drafted with Claude Code, reviewed and edited by the author.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
acinader added a commit to acinader/hledger that referenced this pull request Sep 20, 2026
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring hledgerorg@8ddb03f changes into hledgerorg#2736

should bring hledgerorg#2736 upto parity with hledgerorg#2735

AI usage: drafted with Claude Code, reviewed and edited by the author.
@acinader
acinader force-pushed the i18n-poc branch 2 times, most recently from b0251c0 to 5ef3a4e Compare September 22, 2026 17:18
acinader added a commit to acinader/hledger that referenced this pull request Sep 22, 2026
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring hledgerorg@8ddb03f changes into hledgerorg#2736

should bring hledgerorg#2736 upto parity with hledgerorg#2735

AI usage: drafted with Claude Code, reviewed and edited by the author.
@acinader

Copy link
Copy Markdown
Contributor Author

5ef3a4e is a good example to contemplate.

  1. web: balance report page (#2242) #2739 added some new text that needs to be translated. I have translated it using claude, but I have no way to actually verify.

  2. This pull request adds Internationalization (i18n) capability, and it also adds a mostly "machine translated" de po.

My current take is that we should merge in the i18n stuff, and maybe include the german po file as a starter? But I wouldn't claim that we have a german localization (l10n)?

@acinader
acinader marked this pull request as ready for review September 22, 2026 17:22
@simonmichael simonmichael added A-BUG Something wrong, confusing or sub-standard in the software, docs, or user experience. i18n Internationalisation/localisation-related. and removed A-BUG Something wrong, confusing or sub-standard in the software, docs, or user experience. labels Sep 29, 2026
@simonmichael

Copy link
Copy Markdown
Member

This is very comprehensive and impressive. Great work. We'll go with this for now.

Some changes/discussion needed:

  • Let's follow the recent format of AI usage lines, with approximate output tokens estimates
  • Let's drop the Claude-Session: URL lines, too transient (they all give me session not found).
  • Everything needs to remain human-maintainable; our AI policy says we won't become dependent on AI.
  • Likewise, all code needs to be reviewed and understood by humans (at least to a level of detail where we are confident in it).
    (I18n.hs looks ok at first glance.)

Bikeshedding:

  • "tr translations_" everywhere is verbose - something shorter ? "tr msgs_" ? Or always define tr = I18n.tr translations_ helpers ?

Questions:

  • When do you use i18n rather than tr ?

@simonmichael simonmichael added needs-discussion To unblock: needs more discussion/review/exploration needs-review To unblock: needs more code/docs/design review by someone needs-changes To unblock: needs some changes made, in line with recommendations needs-rebase To unblock: needs to be rebased against latest master branch docs Documentation-related. labels Sep 29, 2026
@simonmichael

Copy link
Copy Markdown
Member

Using the english text as the keys is the the most conventional gettext approach I think. Which is generally a good thing. And it seems convenient to start with. But there will be churn when we want to tweak the ui messages. Is using a sum type of message ids, as in Henning's #2735, compatible with gettext tools ?

@simonmichael

simonmichael commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Let's focus on just the library and CLI reports to begin with.

@simonmichael

simonmichael commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Also let's defer the plural support, if not yet used (?)

But there will be churn when we want to tweak the ui messages.

Ah, but that's a good thing apparently. It signals staleness and updates needed. My claude session recommends combining a sum type for compiler checking plus english text for best translation experience:

Use a Haskell sum type for message ids, as in 2735, and make its English rendering the msgid, as in 2736. The PO files stay conventional and English-keyed. What you gain:

  • Type safety. A typo or a removed message becomes a compile error. That fixes the weakest part of 2736.
  • Simpler extraction. Deriving Enum and Bounded lets a small Haskell program list every message. The regex extractor, and its every-call-on-one-line rule, can go.
  • No contexts. Two constructors can share English text without msgctxt, as in Henning's Total and RightTotal. Each constructor would still produce its own entry, so the catalog would need a context or the constructor name to keep them apart.
  • Parameters fit. A constructor with fields, like one carrying an account name, renders to the English template plus its placeholder values.

The cost is one extra step per new string: add a constructor as well as the literal.

@simonmichael

simonmichael commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Then again: just writing tr before string literals in our code is a very nice gentle first step. And we don't have a huge amount of text and haven't yet experienced the problems compiler checked ids would solve. I'm fine with the current approach for now.

The journal and register handlers still set their titles with plain
setTitle, so a French or German viewer got "journal - hledger-web" in
the tab while the rest of the page was translated. The edit and upload
pages already used setTitleI. Both titles are now catalog entries, with
a translator note, and the German catalog has them. The browser test
for German checks the tab titles too; since the English ones are lower
case, the check cannot pass on an untranslated title.

AI usage: Claude Fable 5.1, ~1k output tokens.
So that the cost of catalogs can be kept an eye on as they grow, --debug
now reports, once, how long parsing the built-in catalogs took, and for
each language loaded: where its catalog came from (built-in and/or the
path of the user's file), how many entries it has, roughly how much text
it holds, and how long loading and merging took. The size is measured
after the timer stops, and none of it is computed without --debug.

  translations: parsed 1 built-in catalogs, 1.8 ms
  translations: loaded de (built-in, ~/.config/hledger/locale/de.po): 179 entries, ~8 KB of text, 0.5 ms

The built-in parse is timed in availableLanguages, because listing the
built-in languages is what forces it.

AI usage: Claude Fable 5.1, ~2k output tokens.
quotedString collected every character of every quoted string into a
String one at a time, then packed it; since nearly all of a catalog is
quoted strings, that was most of the parse. It now takes runs of
ordinary characters as slices of the input and handles only escapes a
character at a time.

A 920 KB catalog of 4000 entries parses in 36 ms instead of 137 ms
(25 MB/s instead of 7), with a quarter of the transient allocation.
The built-in German catalog goes from about 1.8 ms to about 0.9 ms.

AI usage: Claude Fable 5.1, ~3k output tokens.
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring hledgerorg@8ddb03f changes into hledgerorg#2736

should bring hledgerorg#2736 upto parity with hledgerorg#2735

AI usage: Claude Fable 5.1, ~5k output tokens.
Add translations for the balance page

AI usage: Claude Fable 5.1, ~6k output tokens.
acinader added a commit to acinader/hledger that referenced this pull request Sep 29, 2026
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring hledgerorg@8ddb03f changes into hledgerorg#2736

should bring hledgerorg#2736 upto parity with hledgerorg#2735

AI usage: Claude Fable 5.1, ~5k output tokens.
@simonmichael simonmichael removed needs-discussion To unblock: needs more discussion/review/exploration needs-rebase To unblock: needs to be rebased against latest master branch needs-review To unblock: needs more code/docs/design review by someone needs-changes To unblock: needs some changes made, in line with recommendations labels Sep 29, 2026
@simonmichael

Copy link
Copy Markdown
Member

Thanks for the updates. I'm going to merge this as-is if you don't object. I'm feeling the need to move things along, to possibly get preview 5 out on schedule (october 1st).

@acinader

acinader commented Sep 29, 2026 via email

Copy link
Copy Markdown
Contributor Author

@simonmichael

Copy link
Copy Markdown
Member

Or would you rather get the web prs merged first ?

The report renderers now bind tr, trc and trf to the report options once,
in a where clause, and import Hledger.Utils.I18n qualified, so call sites
read `tr "Balance Sheet"` rather than `tr translations_ "Balance Sheet"`
(suggested in the hledgerorg#2736 review). The module docs now say when to mark a
literal with i18n instead of translating it with tr: only for text stored
in data, such as a report spec, before any catalog is at hand.

AI usage: Claude Fable 5.1, ~2k output tokens.
@acinader

Copy link
Copy Markdown
Contributor Author

AI usage. Done: every commit now carries the AI usage line with
a token estimate, and the session URLs are gone. The estimates are for the
tokens that produced each commit's code and text.

I have read I18n.hs; the note in the lib commit saying otherwise predated my
review. I wanted to post the PR before I had a chance to fully review I18n.hs so
we could compare to alternate implementations. I had crossed out that
disclaimer in the PR body, and have now removed it from the commit message too.

My larger point is simply that gettext is a well understood interface that
has many reference implementations and lends itself well to AI support.

Plurals. I'd rather keep trn. The Plural-Forms header is part of what
every translator's tool writes into a catalog. My strong hunch is that anyone
trying to do a full translation will require it, and our i18n machinery will be
obviously incomplete without it.

Naming. The CLI renderers now bind tr, trc, and trf to the report
options once, in a where clause, and import the module qualified, so call sites
read tr "Balance Sheet".

i18n versus tr. tr wherever the text is displayed and a Translations
value is at hand, which is nearly everywhere. i18n only marks a literal
that is stored earlier in data, before any catalog exists: the subreport
titles in the bs/is/cf specs and hledger-ui's menu items. It is the
identity; its job is to make the extractor collect the literal, which tr
then translates at display time. The module docs now say this.

My overall conclusion is that this PR puts the machinery in place at low cost
for us. Someone like @thielema can use the machinery to
do specific, incomplete translations without having to use sed on completed
reports.

The implementation is also complete enough that if someone wanted to do a
comprehensive translation, all the pieces are there. The code changes a full
translation will turn out to need can only be discovered by doing one.

Finally, I'd like this to go in before my other PRs, as it will be easier on me
to rebase the small ones than this big one.

@simonmichael

Copy link
Copy Markdown
Member

Excellent. Thank you very much!

@simonmichael
simonmichael merged commit 24d0a0c into hledgerorg:main Sep 29, 2026
2 checks passed
simonmichael pushed a commit that referenced this pull request Sep 29, 2026
"Monthly Balance Sheet" was an interval word plus a report name, joined
by a space in the code. One adjective form had to fit every report, and
the space could not be dropped. Now each compound report lists its own
title per interval, so a translator renders the whole phrase as the
language needs, with no grammar logic in Haskell. The German catalog
grows from 14 to 44 title entries. English output is unchanged.

this should bring 8ddb03f changes into #2736

should bring #2736 upto parity with #2735

AI usage: Claude Fable 5.1, ~5k output tokens.
@acinader
acinader deleted the i18n-poc branch September 29, 2026 18:25
@simonmichael

Copy link
Copy Markdown
Member

I'm having trouble merging the web PRs, and I see why: the line numbers in the pot file make it very conflicty. This seems like it will be a problem. Could you change the tool to omit the line numbers - not quite as convenient for translators, but I think it's needed.

@simonmichael

Copy link
Copy Markdown
Member

Slightly related, I noticed that the texts from all --lang-supporting apps based on hledger-lib are now centralised in hledger-lib, ie hledger-lib has baked-in knowledge about things that depend on it, which is not good for modularity. But perhaps it is the pragmatic choice.

acinader added a commit to acinader/hledger that referenced this pull request Sep 30, 2026
The template's #: lines now name only the source files a text appears in.

With line numbers, any change that shifted code rewrote hundreds of references
on the next regeneration, so every branch that regenerated the template
conflicted with every other (hledgerorg#2736 review). Now the template changes only when
strings are added, removed, or move to another file.

AI usage: Claude Opus 5.5, ~1k output tokens.
@acinader

acinader commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

Well, that didn't go as smoothly as I was hoping!

I've rebased and fiddled with the open web PRs so you should now be able to merge
them all without causing any new conflicts.

#2755 now writes the template's references as file names only, no line numbers,
like xgettext's --add-location=file, so the template changes only when a
string is added, removed, or moves to another file.

Catalogs in hledger-lib: You're right that it inverts the dependency. My
thinking is one catalog per language is easier for a translator. Ideally we can keep the
single catalog for now and split it if it starts to hurt? I'm also happy to do it now if you prefer,
as I may not understand the full implications.

@simonmichael

simonmichael commented Sep 30, 2026 via email

Copy link
Copy Markdown
Member

simonmichael pushed a commit that referenced this pull request Sep 30, 2026
The template's #: lines now name only the source files a text appears in.

With line numbers, any change that shifted code rewrote hundreds of references
on the next regeneration, so every branch that regenerated the template
conflicted with every other (#2736 review). Now the template changes only when
strings are added, removed, or move to another file.

AI usage: Claude Opus 5.5, ~1k output tokens.
@simonmichael

Copy link
Copy Markdown
Member

It seems to me that --lang is not affecting hledger-web.

@acinader

acinader commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

You're browser is almost certainly sending Accept-Language: en.

I did not make --lang "win" by default and force all responses to be a particular locale.

?_LANG=xx -> cookie -> accept header -> --lang.

hledger-web --serve --lang de -f examples/sample.journal
curl -s http://127.0.0.1:5000/journal | grep -o 'lang="[a-z]*"'
curl -s -H 'Accept-Language: en' http://127.0.0.1:5000/journal | grep -o 'lang="[a-z]*"'

http://127.0.0.1:5000/journal?_LANG=de switches to Derman and sets a cookie to remembers the choice.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs Documentation-related. i18n Internationalisation/localisation-related.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants