Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,17 @@ first release is 3.0.0 because exit codes and the cookie store changed in ways a

## Unreleased

### Added
- `mam/accept_unsnatched` (default off): when the MAM file-name searches return rows but none marked
`my_snatched`, accept the one row that has the release's file type, one of its authors and the same title (same
numbers, subtitle and production: no sibling volume, part, box set or dramatisation); two or more are ambiguous
and none is used. Seen 2026-10-07 with *Before She Knew Him* (Peter Swanson) and every other recent MAM-pass
miss: both file-name searches returned the release's own torrent, but booktree ran minutes after the download
and MAM had not marked it `my_snatched` yet, so the release was left unmatched in the MAM pass. No extra MAM
request: the rows come from the searches booktree already sends, whose text and cache keys are unchanged.
stdout: `No snatched MAM match; using the only unsnatched one that passes the checks: ...`; JSON log:
`mam_attempt: "unsnatched"`. See CONFIG.md, Unsnatched MAM matches.

## 3.0.5 - 2026-10-06

### Fixed
Expand Down
42 changes: 41 additions & 1 deletion CONFIG.md
Original file line number Diff line number Diff line change
Expand Up @@ -64,6 +64,7 @@ A copy of default_config.cfg can be found under the /templates folder. It is re
| cache/mam_empty_hours | | How long an empty MAM answer (or one without a snatched entry) is reused | 24 |
| mam/min_interval_seconds | | Minimum spacing between HTTP requests to MAM (cache hits are free) | 6 |
| mam/max_queries_per_run | | Runaway guard: maximum MAM searches in one run (one container invocation); further searches are skipped with a message and those releases stay unmatched until the next run (0 = unlimited). The 6-second spacing is the real safety net; this only stops a loop | 3000 |
| mam/accept_unsnatched | | 1 = when the MAM file-name searches return rows but none marked as snatched by you, accept the one row that passes the file type, author and title checks (no extra MAM request; see [Unsnatched MAM matches](#unsnatched-mam-matches-mamaccept_unsnatched-default-off)) | 0 |
| json_log | | Path of a JSON-lines run log, or `true` for `booktree_log_<timestamp>.jsonl` next to the CSV; same as `--json-log` | |
| refresh | | List of releases (name or path) to re-process ignoring cached answers and the processed marker; same as `--refresh` | |
| pins | | List of `RELEASE=ASIN` strings: use that Audible ASIN for the release and re-process it now; same as `--pin` (see [Correcting a match](#correcting-a-match)) | |
Expand Down Expand Up @@ -320,6 +321,45 @@ apart, and afterwards once per TTL.
its "already processed" marker are ignored while everything else is served from cache. `--no-cache` still does this
for the whole run.

### Unsnatched MAM matches (`mam/accept_unsnatched`, default off)

The MAM pass searches MAM's file-name index for the release's file name (with its authors, then without) and keeps
only torrents MAM marks as snatched by you (`my_snatched`), so the match is the torrent you downloaded. booktree
usually runs minutes after a download completes (a download-complete hook), and MAM sets that mark minutes to hours
later: the search finds the release's own torrent, not yet marked, and the release stays unmatched in the MAM pass
(seen with every MAM-pass miss in October 2026, such as *Before She Knew Him* by Peter Swanson).

With `"mam": {"accept_unsnatched": 1}`, when neither file-name search returned a snatched row but they did return
rows, booktree looks at those same rows (no extra MAM request; same cache) and accepts one when it is the only row,
by MAM id, that passes all of these checks:

* the release's file type (`m4b`, `mp3`, ...) is among the row's file types;
* one of the release's authors (id3, or parsed from the release name where the tags are junk) is one of the row's
authors, by the same name comparison used for snatched rows;
* the same title: main titles with a token-sort ratio of 90 or more and the same numbers (also written out, as roman
numerals or as "Book N"); the same production (abridged, dramatised, GraphicAudio, full cast); matching subtitles
when both have one; no subtitle that only the release has; and a subtitle that only MAM has only when it is the
name of that torrent's own series and names no bundle or part;
* the same series position: a "Book N" written in one title must appear in the other title or among the other
side's series parts (MAM's `series_info` for the row, the release's own series tags), otherwise the row is refused.

So *The Viscount and the Witch* is not *The Witch*, *Cradle: Soulsmith* is not *Cradle: Unsouled*, *Part 1* is not
*Part 2*, *Volume II* is not *Volume I*, *Dune: The Complete Saga* and *Dune: Part One* are not *Dune*, *Mistborn:
Secret History* is not *Mistborn*, *Mother of Learning, Book 2* is not a row *Mother of Learning* that MAM lists as
part 3, while *Leviathan Wakes* (tagged The Expanse #1) is the row *Leviathan Wakes: The Expanse, Book 1* in the
series The Expanse.

Two or more passing rows are ambiguous and none is used; a snatched row is always used as before. MAM rows carry no
runtime, so these checks are the only protection: in `mam-audible` the Audible search that follows is built from
the MAM match's title and authors, and the Audible runtime there is a preference, not a rejection, so a wrong MAM
match would carry through. Nothing is checked with `flags/verbose` 0 and `flags/ebooks` 0, where booktree does
not rank MAM results at all. stdout says `No snatched MAM match; using the only unsnatched one that passes the checks: <title> by
<authors>`; the JSON log marks such a book with `mam_attempt: "unsnatched"` (and `match.attempt` is
`mam-unsnatched` when the MAM record is the one filed).

It is off by default because it accepts a MAM torrent not marked as yours, judged by file type, author and title
alone. Turn it on for the config that uses `mam` or `mam-audible` if booktree runs right after downloads complete.

## JSON run log (`--json-log [PATH]`)

In addition to the CSV, one JSON object per line: a `book` record per processed release and a final `run` record.
Expand All @@ -344,7 +384,7 @@ The CSV columns, file name and stdout are unchanged; the JSON log is the structu
~~~

`match.attempt` says what produced the match: `pinned`, `candidates`, `parsed`, `parsed-authors`, `swapped`,
`title-only`, `legacy` (upstream's own search), `mam`, or `log` (taken from the input log in `log` mode).
`title-only`, `legacy` (upstream's own search), `mam`, `mam-unsnatched` (the MAM match is a torrent not yet marked snatched, `mam/accept_unsnatched`; see also `mam_attempt`), or `log` (taken from the input log in `log` mode).
`queries[].cache_key` is the hash printed in `Checking cache: <kind>/<hash>`; a query that failed carries `error`,
one skipped by the MAM budget carries `skipped`. `id3.duration_s` is the first file's duration while
`expected_duration_min` is the whole release. One `run` record is written per `paths` entry. Give the path as
Expand Down
190 changes: 186 additions & 4 deletions myx_classes.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,11 +16,140 @@
import myx_hints
import myx_names
import copy
import html
from thefuzz import fuzz

#Module variables
authMode="login"
verbose=False

#Accepting a MAM torrent not marked my_snatched (Config/mam/accept_unsnatched, MAMBook.pickUnsnatched)
MAM_TITLE_MIN = 90 # token-sort ratio for the main titles, and for the subtitles when both have one
MAM_MAX_TITLE = 200 # untrusted titles (id3, MAM rows) are cut before the title regexes run
_SUBTITLE_SPLIT = re.compile(r"\s*:\s*|\s+[-\u2013\u2014]\s+")
_SERIES_POSITION = re.compile(r"\bbook\s*\d+(?:\.\d+)?\b", re.IGNORECASE)
# a subtitle that only MAM has and that names a bundle or a part, not this book ("Dune: The Complete Saga",
# "Dune: Part One"); checked on the raw subtitle, before noise words such as "complete" are removed
_BUNDLE_WORDS = re.compile(r"\b(?:complete|saga|collection|box\s*set|boxed|omnibus|trilogy|series|books|volumes?|"
r"vol|bundle|part|chapter|episode|season)\b", re.IGNORECASE)
# numbers that tell volumes apart, written out or as roman numerals ("Volume II", "Part One")
_NUMBER_WORDS = {w: str(n) for n, w in enumerate(
"zero one two three four five six seven eight nine ten eleven twelve thirteen fourteen fifteen sixteen seventeen "
"eighteen nineteen twenty".split())}
_NUMBER_WORDS.update({w: str(n) for n, w in enumerate(
"_ i ii iii iv v vi vii viii ix x xi xii xiii xiv xv xvi xvii xviii xix xx".split()) if n})
_EDITION = re.compile(r"(?<!un)\babridged\b|\bdramati[sz]ed\b|\bdramati[sz]ation\b|\bgraphic\s?audio\b|\bfull\s?cast\b",
re.IGNORECASE)
_NUMBER_WORDS.update({"first": "1", "second": "2", "third": "3", "fourth": "4", "fifth": "5"})


def _titleWords(text):
"""Lower-case words for comparing titles: HTML entities decoded (MAM rows carry &#039; and &amp;), accents and
apostrophes removed ("Ender's" = "Enders", "Shōgun" = "Shogun"), noise words (Unabridged, m4b, ...) and other
punctuation removed."""
t = html.unescape(str(text or "")).replace("&", " and ")
t = myx_utilities.strip_accents(re.sub(r"['\u2019]", "", t))
t = myx_names.NOISE_WORD.sub(" ", t)
return " ".join(re.sub(r"[^\w\s]|_", " ", t).lower().split())


def _numbers(words):
"""The numbers in a _titleWords string: digits, number words and roman numerals ("i" counts: "Volume I")."""
return sorted(w if w.isdigit() else _NUMBER_WORDS[w] for w in words.split() if w.isdigit() or w in _NUMBER_WORDS)


def _clip(title):
"""html-unescaped, whitespace collapsed and cut to 2 x MAM_MAX_TITLE: id3 tags and MAM rows are untrusted,
and the title regexes are only linear on bounded, collapsed input."""
return " ".join(html.unescape(str(title or ""))[:MAM_MAX_TITLE * 4].split())[:MAM_MAX_TITLE * 2]


def _edition(title):
"""Markers of a different production of the same title: abridged (not unabridged), dramatised, GraphicAudio."""
return sorted(set(m.lower().replace(" ", "")[:6] for m in _EDITION.findall(title)))


def _titleParts(title):
"""(main title, subtitle, raw subtitle): the first two as _titleWords with "Book N" series positions dropped and a
"a novel" style subtitle emptied; the raw subtitle only html-unescaped."""
t = _clip(title)
t = re.sub(r"\s*\((?:un)?abridged\)", "", t, flags=re.IGNORECASE)
parts = _SUBTITLE_SPLIT.split(t, maxsplit=1)
raw = parts[1] if len(parts) > 1 else ""
main = _titleWords(_SERIES_POSITION.sub(" ", parts[0]))
sub = _titleWords(_SERIES_POSITION.sub(" ", raw))
if myx_names.SUBTITLE_NOISE.match(sub):
sub, raw = "", ""
return main, sub, raw


def sameMamTitle(ours, theirs, series=()):
"""The title check for an unsnatched MAM row (pickUnsnatched): the same book, not a sibling volume, a box set or another part.
Main titles must agree (token-sort ratio >= MAM_TITLE_MIN) and carry the same numbers, also written out
or in roman numerals ("Part 1" is not "Part 2", "Volume II" is not "Volume I"); when both have a subtitle those
must agree too ("Cradle: Unsouled" is not "Cradle: Soulsmith"); a subtitle only we have is refused ("Thrawn:
Treason" is not "Thrawn"); one only MAM has is accepted only when it is the name of the candidate's own series
(`series`, from MAM's series_info) and names no bundle or part: "Leviathan Wakes: The Expanse, Book 1" in series
The Expanse yes; "Dune: The Complete Saga", "Dune: Part One", "Mistborn: Secret History" no."""
om, osub, _ = _titleParts(ours)
tm, tsub, traw = _titleParts(theirs)
if not om or not tm or fuzz.token_sort_ratio(om, tm) < MAM_TITLE_MIN:
return False
# "Book 2" is not "Book 3" (the positions are dropped from the words compared below), and an abridged, dramatised
# or GraphicAudio production is not the plain one
position = seriesPosition
if position(ours) != position(theirs) and position(ours) and position(theirs):
return False
if _edition(_clip(ours)) != _edition(_clip(theirs)):
return False
if _numbers(f"{om} {osub}") != _numbers(f"{tm} {tsub}"):
return False
if osub and tsub:
return fuzz.token_sort_ratio(osub, tsub) >= MAM_TITLE_MIN
if osub:
return False
if _BUNDLE_WORDS.search(traw): # also when noise-word removal left nothing ("Dune: Complete")
return False
if not tsub:
return True
names = [_titleWords(_clip(n)) for n in series or ()]
return any(n and fuzz.token_set_ratio(n, tsub) >= MAM_TITLE_MIN for n in names)


def _num(part):
"""A series part as a comparable string: "01", "1.0" and 1 are all "1"; "" when there is none."""
try:
return format(float(str(part).strip()), "g") if str(part).strip() else ""
except ValueError:
return str(part).strip().lower()


def seriesPosition(title):
"""The "Book N" numbers written in a title, as a sorted list of _num strings."""
return sorted(_num(n) for n in re.findall(r"\d+(?:\.\d+)?", " ".join(_SERIES_POSITION.findall(_clip(title)))))


def samePosition(ours, theirs):
"""Series position check for an unsnatched MAM row (pickUnsnatched), between two Books: a "Book N" written in one
title must be matched by the other's title or, failing that, by one of its series parts; when the other side has
no position at all, the row is refused ("Mother of Learning, Book 2" is not a row "Mother of Learning" that MAM
files as part 3, nor one with no part)."""
def positions(b):
return set(seriesPosition(b.title)), {_num(s.part) for s in b.series if _num(s.part)}
ot, op = positions(ours)
tt, tp = positions(theirs)
if ot and not tt:
return bool(ot & tp)
if tt and not ot:
return bool(tt & op)
return True # both titles carry positions (sameMamTitle compared them) or neither does


def _oneLine(text):
"""An untrusted field (id3 tag, MAM row) for a one-line message: whitespace and control characters collapsed."""
return " ".join("".join(c if c.isprintable() else " " for c in str(text or "")).split())[:MAM_MAX_TITLE]


#Author and Narrator Classes
@dataclass
class Contributor:
Expand Down Expand Up @@ -483,6 +612,7 @@ class MAMBook:
parsedName:dict=None
refresh:bool=False
matchAttempt:str=""
mamAttempt:str=""

def getRunTimeLength(self):
#add all the duration of the files in the book, and convert into minutes
Expand Down Expand Up @@ -1105,14 +1235,25 @@ def getMAMBooks(self, cfg, bookFile:BookFile):

# Search using book key and authors (using or search in case the metadata is bad)
print(f"Searching MAM for\n\tTitleFilename: {title}\n\tauthors:{authors}")
books=myx_mam.getMAMBook(cfg, titleFilename=title, authors=authors, extension=extension, refresh=self.refresh)
# Config/mam/accept_unsnatched: also keep the rows not marked my_snatched (no extra request; pickUnsnatched)
pool = [] if myx_mam.acceptUnsnatched(cfg) else None
keep = {} if pool is None else {"unsnatched": pool}
books=myx_mam.getMAMBook(cfg, titleFilename=title, authors=authors, extension=extension, refresh=self.refresh, **keep)

# was the author inaccurate? (Maybe it was LastName, FirstName or accented)
# print (f"Trying again because Filename, Author = {len(self.mamMatches)}")
if len(books) == 0:
#try again, without author this time
print(f"Widening MAM search using just\n\tTitleFilename: {title}")
books=myx_mam.getMAMBook(cfg, titleFilename=title, extension=extension, refresh=self.refresh)
books=myx_mam.getMAMBook(cfg, titleFilename=title, extension=extension, refresh=self.refresh, **keep)

# neither file-name search found a snatched torrent, but they did return rows: accept the one that passes the
# checks, if there is exactly one (pickUnsnatched); the ranking below then treats it like a snatched one
# (not when the ranking below does not run: with verbose off it never does, an upstream quirk)
fromUnsnatched = False
if len(books) == 0 and pool and (verbose or ebooks):
books = self.pickUnsnatched(cfg, rankBook if rankBook is not None else self.ffprobeBook, bookFile, pool)
fromUnsnatched = bool(books)

#Find the best match
self.mamMatches = books
Expand Down Expand Up @@ -1156,8 +1297,8 @@ def getMAMBooks(self, cfg, bookFile:BookFile):
targetBook = '|'.join([book.title, book.getAuthors(), book.getSeriesParts()])

for abook in books:
#if this book is snatched, include in the match
if abook.snatched:
#if this book is snatched, include in the match (an accepted unsnatched one passed pickUnsnatched)
if abook.snatched or fromUnsnatched:
#the author is known, check if this book is this authors book
#otherwise, if maybe this title is close enough
#print (f"{abook.title} by {abook.authors}...")
Expand Down Expand Up @@ -1189,12 +1330,53 @@ def getMAMBooks(self, cfg, bookFile:BookFile):
self.bestMAMMatch = books[0]


if fromUnsnatched and self.bestMAMMatch is not None:
self.matchAttempt = "mam-unsnatched"
self.mamAttempt = "unsnatched"

#pprint(self.bestMAMMatch)
if (books is not None):
return self.bestMAMMatch
else:
return None

def pickUnsnatched(self, cfg, book, bookFile, pool):
"""Config/mam/accept_unsnatched: the file-name searches found rows but none marked my_snatched. That is usually
the release's own torrent: the hook runs booktree minutes after a download and MAM sets my_snatched later.
`pool` holds those rows as (key, Book, file types), from both searches. A row is accepted only when it is the
one row (by MAM id) that has the release's file type, one of its authors (exact, as for snatched rows), the same
title (sameMamTitle: no sibling volume, part, bundle or other production) and the same series position
(samePosition: a "Book N" in one title against the other's title or series part); two or more are ambiguous and
none is used. MAM rows carry no runtime. Returns [the Book] or []."""
if book is None:
return []
parsed = self.getParsedName(book, cfg) or {}
title = _clip(book.title)
# a title that only repeats the release name is junk, unless the release name is just the title
junk = myx_names.isJunkTitle(title) or (myx_names.isJunkTitle(title, parsed.get("source") or self.name) and
_titleWords(parsed.get("title") or "") != _titleWords(title))
if junk or not book.authors or myx_names.isJunkAuthors(book.authors):
print("No snatched MAM match; no usable title and author to check the unsnatched ones")
return []
ext = bookFile.getExtension().lower()
rows = {}
for key, abook, types in pool:
rows.setdefault(key, (abook, types))
# the silent checks first: isThisMyAuthorsBook prints the row's title (verbose), and a row that is not ours
# need not reach stdout
passed = [abook for abook, types in rows.values()
if ext and ext in types and sameMamTitle(title, abook.title, [x.name for x in abook.series])
and samePosition(book, abook) and myx_utilities.isThisMyAuthorsBook(book.authors, abook, cfg)]
if len(passed) == 1:
print(f"No snatched MAM match; using the only unsnatched one that passes the checks: "
f"{_oneLine(passed[0].title)} by {_oneLine(passed[0].getAuthors())}")
return passed
if passed:
print(f"No snatched MAM match; {len(passed)} unsnatched ones pass the checks: ambiguous, not using any")
else:
print(f"No snatched MAM match; no unsnatched one passes the title, author and file type check ({len(rows)} row(s))")
return []

def getHashKey(self):
return myx_utilities.getHash(self.name)

Expand Down
Loading
Loading