Skip to content

Group imports into one drive with a document and native issues table - #21

Merged
michielbdejong merged 1 commit into
mainfrom
claude/reflector-rs-issue-18-wjg3wb
Sep 7, 2026
Merged

michielbdejong merged 1 commit into
mainfrom
claude/reflector-rs-issue-18-wjg3wb

Conversation

@michielbdejong

Copy link
Copy Markdown
Contributor

Closes #18.

Summary

  • One drive per imported dataset, not per record namespace. AtomicStorage previously keyed a Drive off each record's own namespace, so a nested collection like issue comments (namespaced per parent issue, e.g. owner/repo/42) fragmented into a separate Drive per issue instead of sharing the repository's one Drive. AtomicStorage now takes a dataset (with_dataset, e.g. owner/repo) — computed by a new Config::dataset_namespace() from API_CONSTANTS' values — and groups every record it ever writes, root or nested, under that one Drive.
  • A DocumentV2 and a native AD Table per root-level resource. Created lazily (idempotent — has_stored_resource-guarded like the existing ensure_drive) the first time a matching root-level record is put. A root-level record's parent becomes its table (so it renders as a row); nested records (comments) keep parent pointing straight at the drive — same drive either way. drive always points at the drive itself, independent of parent, matching how atomic_lib's own rights resolution (parent chain) and fan-out (drive pointer) are documented to work.
  • Saved-drive registration. When DRIVE_OWNER names an agent that already has a local resource, the drive is added to that agent's drives property (idempotently, on every sync) — a generalized, Storelike-level version of atomic_lib's own (private, Db-specific) push_drive_to_list.
  • Migration. A drive left over from the old per-namespace scheme, under this same dataset, is removed on the next sync. The records that pointed at it get re-pointed to the real drive automatically the next time they're put (persist always rewrites parent/drive from scratch, and every sync re-puts every record today).
  • Source URL column. Added html_url to the vendored issue schema so it flows through the existing (generic, GitHub-agnostic) ontology derivation as a table column alongside number/title/state/body/author/updated-at, with no syncables-rs changes.

Known scope limits (called out for review)

  • documentContent is a plain Markdown property, not the data browser's own collaborative loro-prosemirror CRDT tree. I looked for a way to construct that tree from Rust (see the session log) and found no reference implementation in this codebase or a vendored crate for it — hand-rolling an undocumented binary CRDT shape risked producing a document that fails to open in the real editor, which seemed worse than a plain, valid Markdown body. The Table is still structurally parented under the Document either way.
  • The table reuses the existing per-resource ontology Class as its classtype rather than a second, narrower Class built just for the table. Since every property a Class recommends/requires becomes a column, the issues table ends up with a few extra columns (id, state_reason, labels) beyond the seven the issue names. Duplicating those property subjects into a parallel row Class felt like more risk (two class definitions to keep in sync) than benefit for a few extra columns.
  • Not manually verified in the app. This environment has no atomic-server/data-browser frontend to click through — the resource shapes are checked against ontola/atomic-server's own Rust (hierarchy.rs, resources.rs) and TypeScript (createTableFromSpec.ts, useTableData.ts) source at the pinned revision, and against unit tests here, not a running UI.

Testing

  • cargo fmt -- --check
  • cargo test (43 tests, incl. new coverage for containment record→table→document→drive, nested collections staying off the table, saved-drive registration idempotence, and legacy-drive migration)
  • cargo clippy --all-targets -- -D warnings
  • cargo build
  • cargo build --target wasm32-unknown-unknown --lib
  • cargo clippy --target wasm32-unknown-unknown --lib -- -D warnings

All pass locally. Docs (README.md, CLAUDE.md, .env.example) updated alongside the code, and a session log added under docs/ai-logs/sessions/ per this repo's NLnet disclosure policy.

🤖 Generated with Claude Code

https://claude.ai/code/session_01RDCvJurQwKTv2STsa5Jfqc


Generated by Claude Code

…18)

A Drive used to be keyed off each record's own namespace, so a nested
collection like issue comments (namespaced per parent issue) fragmented
into a separate Drive per issue instead of sharing the repository's one
Drive. AtomicStorage now groups every record — root and nested — under
one Drive per imported dataset (`with_dataset`, derived from
API_CONSTANTS), with a DocumentV2 and a native AD table per root-level
resource under it: root records file under the table (so they render as
rows), nested records keep filing straight under the drive. The drive is
also registered on DRIVE_OWNER's saved-drive list, and a leftover
per-namespace drive from before this fix is migrated away automatically.

Also adds `html_url` to the vendored issue schema as its source URL, so
it flows through as a table column alongside number/title/state/body/
author/updated-at with no syncables-rs changes.

Claude-Session: https://claude.ai/code/session_01RDCvJurQwKTv2STsa5Jfqc
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@michielbdejong
michielbdejong merged commit 0757108 into main Sep 7, 2026
1 check passed
michielbdejong pushed a commit that referenced this pull request Sep 7, 2026
Merging origin/main (#21, which grouped every record under one Drive/
Document/Table per imported dataset) auto-merged cleanly but left
Config::dataset_namespace() referencing a `constants` field that this
branch had already moved onto each PlatformConfig, and a single
AtomicStorage shared across all configured platforms — which would
collide every platform's records into the same dataset instead of
keeping each platform's own Drive.

Moves dataset_namespace() onto PlatformConfig (derived from that
platform's own constants) and builds one AtomicStorage per platform,
sharing the same underlying Db, matching the one-AtomicStorage-per-
dataset pattern #21 already established for a single run with multiple
namespaces.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015uMtEr3myMSRyBMaSWP3k5
michielbdejong added a commit that referenced this pull request Sep 7, 2026
* Support reflecting more than one platform in a single run

Restructures spec/ into one subfolder per platform and adds a PLATFORMS
environment variable: unset, behavior is unchanged (one `github` platform
from the historical unprefixed variables); set, each named platform reads
its own <PLATFORM>_-prefixed OpenAPI document/overlays/token/constants,
falling back to built-in defaults for `github` and a new `google-calendar`
platform (read-only: calendar-list entries and events). main.rs now syncs
every configured platform into the same store, applying the GitHub OAuth
fallback only to a platform actually named `github`.

Clockify, also requested in the issue, needs syncables-rs to support
apiKey-header credentials first (it only models Authorization: Bearer
today) — documented in the README rather than added without working auth.

Closes #15.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015uMtEr3myMSRyBMaSWP3k5

* Fix dataset_namespace after merging main's per-dataset drive grouping

Merging origin/main (#21, which grouped every record under one Drive/
Document/Table per imported dataset) auto-merged cleanly but left
Config::dataset_namespace() referencing a `constants` field that this
branch had already moved onto each PlatformConfig, and a single
AtomicStorage shared across all configured platforms — which would
collide every platform's records into the same dataset instead of
keeping each platform's own Drive.

Moves dataset_namespace() onto PlatformConfig (derived from that
platform's own constants) and builds one AtomicStorage per platform,
sharing the same underlying Db, matching the one-AtomicStorage-per-
dataset pattern #21 already established for a single run with multiple
namespaces.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015uMtEr3myMSRyBMaSWP3k5

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Create a user-owned AD drive with a document and issues table for imported GitHub data

2 participants