Skip to content

Repository files navigation

atif-sql

CI security CodeQL Scorecard Python 3.13+ License Apache-2.0

ATIF-native analytics over agent trajectories.

Claude Code sessions (~/.claude/projects/**/*.jsonl) are converted to ATIF (Harbor's Agent Trajectory Interchange Format, see the Harbor ATIF RFC: RFC-0001 in https://github.com/laude-institute/harbor), materialized as a corpus, and queried through DuckDB views. Converting once at the boundary — with an explicit, tested fidelity policy for what upstream drops — beats re-deriving trajectory semantics inside every SQL view.

Install

One command, one package, every capability. Conversion, materialization, the DuckDB surface, the LLM analytics pipelines, and semantic search are all in the box — there are no extras to choose and nothing to install afterwards to make a command work.

uv tool install atif-sql     # the CLI on PATH
uvx atif-sql schema          # or run it without installing

Python 3.13 or newer. The install is substantial and deliberately so: 113 runtime dependencies, about 1.15 GiB on disk, because the analytics and vector paths carry polars, pyarrow, scipy, scikit-learn, umap-learn, hdbscan, lancedb, and duckdb. Prebuilt wheels cover CPython 3.13 on manylinux x86_64, macOS arm64, and Windows x86_64; Linux aarch64 compiles hdbscan from source, which needs a C toolchain. Alpine and other musl targets are not supported. RELEASING.md carries the measurements.

atif-sql analyze, atif-sql embed, and atif-sql search call Amazon Bedrock and cost money per invocation. Each one is dry-run by default and spends only when asked. Nothing else in the tool needs a credential.

How the source is organized

These seven directories under packages/ are internal structure, not seven installs. They exist so import-linter can enforce the layer and independence contracts at the source level; the only thing documented as installable is the atif-sql CLI above.

Directory What
atif-converter Harbor ClaudeCode adapter wrapper + fidelity policy (loss accounting per session)
atif-corpus Corpus materialization: discovery, watermarks, quiescence, atomic artifact writes
atif-duck DuckDB views + macros over the materialized corpus (core surface plus the v2 analytics surface)
atif-models Model alias registry + structured-output LLM client; no other package hardcodes a model id
atif-analytics eight v2 pipelines — five LLM (classify, trajectory, conflicts, friction, perceived) and three structural (cluster, terms, community)
atif-embed Cohere Embed v4 on Bedrock + LanceDB vector store + embedding backfill
atif-cli the composition root, and the source of the atif-sql command: convert, materialize, status, query, analyze, embed, search, examples, schema, cron

Quick start

Build the corpus from the local transcripts, then query it:

atif-sql materialize                   # discover sessions, convert, write the corpus
atif-sql status                        # corpus freshness, read-only
atif-sql query 'SELECT * FROM sessions LIMIT 5'

For one session at a time, atif-sql convert <session.jsonl> converts and audits it in place.

Working on atif-sql itself is a different setup — a clone, mise, and mise run check as the definition of done. CONTRIBUTING.md has it.

Agent workflow: schema → examples → query

atif-sql schema                    # every view + macro signature (<50 ms)
atif-sql examples                  # tested example queries, grouped core/analytics/vss
atif-sql query 'SELECT * FROM tool_rank(30) LIMIT 10'

atif-sql examples (or atif-sql query --examples) emits runnable queries derived from the catalog — not hardcoded strings — and every one is executed by the test suite against a fixture corpus, so the listing cannot rot. Piped output is JSON; filter with --requires core|analytics|vss and --category view|table-macro|scalar-macro.

Contributing

Setup, the five gates mise run check runs, the import-linter contracts you will trip, and the Conventional Commit rule the commit-msg hook enforces: CONTRIBUTING.md.

Releases are cut by commitizen and published to PyPI over OIDC Trusted Publishing — RELEASING.md covers the flow, the versioning model, and the install-weight measurements.

Security

Report a vulnerability privately through GitHub's advisory form, not in a public issue. Supported versions, the disclosure expectations, and what counts as a vulnerability in a tool that reads local transcripts are in SECURITY.md.

License

Apache License 2.0. Each of the seven module directories carries the same license file, so a published distribution ships it too.

About

ATIF-native analytics over Claude Code agent trajectories: convert sessions to ATIF, materialize a corpus, query it with DuckDB.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages