Self-hosted AI accountant: aggregates accounts, tracks goals, surfaces recommendations, and graduates from read-only analysis to human-approved execution to bounded autonomy. Full design in ai-finance-platform-spec.md.
Status: Phase 2 complete and running on real data. Phase 3a (approval-gated paper trading) is built and waiting on broker keys. Local-only, zero hosting cost. Plaid production access is approved. Gmail receipt ingestion is live on real mail.
License: PolyForm Noncommercial 1.0.0. Use it, change it, run it for your own household, share it — commercial use is not granted. You may not sell it, resell it, or run it as a paid service. This is not an OSI-approved open-source licence, and that is deliberate.
This is one person's self-hosted build, shared as-is. It assumes a single owner and a local Supabase stack — there is no multi-tenancy anywhere and none is planned. Everything instance-specific lives in
.env; see Environment keys.docs/carries the security and data-retention policies Plaid's questionnaire asks for, as templates with<OWNER NAME>placeholders — fill them in and generate your own PDFs.
No personal data in this repository, ever. A
pre-commithook (scripts/hooks/pre-commit, wired viacore.hooksPath) blocks staged secrets, email addresses, long account/reference numbers, and anything listed in a gitignored.pii-denylist. Examples in comments are invented on purpose. See Keeping personal data out.
- Monorepo:
apps/web(Next.js 14 App Router),apps/worker(node-cron workers),packages/shared(DB client, Plaid client, token crypto, categorization engine). - Database: Supabase local stack (Docker, Postgres 17) with the complete
spec §4 schema — 40+ tables covering accounts, transactions, categorization,
email receipts/anticipation, goals with cost attribution, recaps, the
recommendation queue, agent config, self-improvement, and safety tables.
RLS locks every row to the single allow-listed owner;
audit_logis append-only (enforced by trigger, even against the service role). - Auth: single owner account, password + mandatory TOTP 2FA (enrollment
forced on first login). Public signup disabled at the auth server. The
/accountpage (click your email in the sidebar) has in-app password change (TOTP-gated), the auto-lock cadence (fresh TOTP demanded when the last one is older than the setting, default 1 hour), and Gmail connections. - Network posture: web app binds 127.0.0.1 only; Docker publishes all
Supabase ports on 127.0.0.1 (
scripts/harden-docker-loopback.sh, already applied). Nothing is reachable from the LAN. - Banking (Phase 1): Plaid Link with encrypted access tokens, 6-hour sync of transactions/balances/holdings/liabilities, the categorization pipeline (rules → merchant map → Plaid baseline → review inbox), recurring detection, business layer, and net-worth snapshots.
- Receipts (Phase 2): multi-mailbox Gmail ingestion with per-inbox label mapping, LLM parse fallback, the anticipation engine, scored reconciliation, and the vendor watchlist — see Phase 2 notes.
- The agent (Phase 2/3a): daily analysis run on
claude-sonnet-5producing schema-validated recommendations into the Approval Queue — advisory alerts, andtradeproposals once trading is switched on. - Investments (Phase 3a): broker positions with unrealized P&L and allocation, recent orders, and the platform's own execution attempts — refusals included, since a guardrail that holds looks like nothing happening.
- Reports (Phase 2): cash flow Sankey, MoM/YoY trends, a saved custom report builder, the business tax export, and the recaps reader.
- Goals (Phase 2): three-step wizard with semantic linkage and a historical preview, nightly contribution matching, pace math, and true-cost panels.
- Recaps (Phase 2): weekly + monthly, deterministic math scored and narrated by Claude under a checked no-invented-numbers rule.
- Workers: sync / agent / executor under systemd, heartbeating every 30s.
- Reboot persistence: systemd user units + lingering; everything returns after a reboot.
| Service | URL |
|---|---|
| Web app | http://localhost:3141 |
| Supabase API | http://127.0.0.1:54321 |
| Postgres | postgresql://postgres:postgres@127.0.0.1:54322/postgres |
| Studio | http://127.0.0.1:54323 |
| Mailpit (captured email) | http://127.0.0.1:54324 |
System tray (primary control): a "Life Command" icon lives in the KDE
system tray — green sparkline when the stack is up, grey when down, with live
health in the tooltip. Left-click opens the dashboard (auto-starting the stack
if needed); right-click: Start / Stop (frees ~2 GB RAM) / Restart / Status /
Studio / Quit. Runs as the finance-tray user service
(scripts/tray.py, pure python3-gobject).
App menu entry: "Life Command" is also in the KDE app launcher. Pin it to
the task bar from the app menu (right-click the menu entry → Pin), not from
a running window — Chrome app-mode windows on Wayland identify as "chrome",
so window pins break. For a perfectly pinnable window, open the dashboard in
Chrome and use menu → Cast, save and share → Install page as app — the PWA
manifest gives it a proper identity. Both installed by
scripts/install-desktop.sh; control logic in scripts/stack-ctl.sh.
Everything is a systemd user service and starts at boot. Manual controls:
npm run svc:status # all three services
npm run svc:restart # bounce web + workers
npm run svc:logs # follow logs
npm run svc:stop # stop web + workers (supabase keeps running)After changing supabase/config.toml: systemctl --user restart finance-supabase
(its start also re-syncs env files). After changing web code: npm run build
then npm run svc:restart. Dev loop: npm run dev (Next dev server on 3000 —
pass -p if that clashes) or use the systemd build.
npm install
npx supabase start # pulls Docker images, applies migrations
node scripts/sync-env.mjs # writes .env files from the running stack
OWNER_EMAIL=you@example.com node scripts/create-owner.mjs
npm run build
bash scripts/install-services.sh
npm run svc:start
git config core.hooksPath scripts/hooks # PII guard; see belowThen add your own keys to the root .env (Plaid, Anthropic, Google, ntfy) and
re-run npm run env:sync — it preserves keys you added and rewrites only the
ones derived from the running stack.
The application is portable: Node, Next and the Supabase CLI behave the same
everywhere, and nothing in apps/ or packages/ assumes a POSIX path. What
differs is only the layer that keeps it running at login.
| Runs | Autostart at login | Notes | |
|---|---|---|---|
| Linux | yes | yes — systemd user units | Developed here; the reboot path is tested |
| macOS | yes | yes — launchd agents | Generated by npm run svc:install |
| Windows | yes | manual — Task Scheduler | npm run up works; no service integration yet |
npm run svc:* dispatches to whatever the host actually has, so the same
commands work on Linux and macOS:
npm run svc:install # systemd units, or ~/Library/LaunchAgents plists
npm run svc:start
npm run svc:status
npm run svc:logsEvery platform can run it in the foreground instead, which is fully functional and simply does not survive a reboot:
npm run upDocker Desktop must be running before Supabase starts on macOS and Windows. Set it to launch at login in its own settings — the launchd agent waits for the CLI, not for Docker itself.
Windows has no autostart integration yet. npm run up runs everything;
for login start, create two Task Scheduler tasks with Start in set to this
directory, running node scripts\\launch-web.js and
node scripts\\launch-workers.js. npm run svc:start prints these
instructions rather than failing.
Not portable, Linux-only, and not required: the desktop launcher, the
system-tray controller (scripts/tray.py, KDE/Wayland), and
scripts/harden-docker-loopback.sh. These are conveniences on the author's
machine; nothing depends on them.
The part that takes longest is not code. Sandbox works in minutes; production took roughly two weeks of back-and-forth, almost all of it waiting on reviews.
-
Sign up at dashboard.plaid.com.
-
Team Settings → Keys — copy
client_idand the Sandbox secret. -
Put them in the root
.env, thennpm run env:sync:PLAID_CLIENT_ID=... PLAID_SECRET=... PLAID_ENV=sandbox -
Restart (
npm run svc:restart) and link anything from the app. Sandbox credentials areuser_good/pass_good.
Dashboard → Request production access. You are asked for a company, a product description, and expected volume. A personal, single-user, self-hosted tool is an accepted answer — say exactly that. Approval took a few days.
Plaid then requires a Compliance Center app profile before any production key works:
- App name, description, logo (a 1024×1024 PNG; a plain wordmark is fine).
- A public-facing URL describing the app. Plaid checks this. A single page on a domain you control is enough — it needs to describe what the app does, what data it accesses and why, and link a privacy policy.
- Security questionnaire — encryption at rest and in transit, access
control, retention.
docs/in this repository carries the policy documents that answer it, as templates with<OWNER NAME>placeholders. Fill them in, export to PDF, upload. Do not commit the filled-in PDFs —docs/*.pdfis gitignored for that reason.
Once approved, swap the secret and env:
PLAID_SECRET=<production secret>
PLAID_ENV=production
npm run env:syncis not optional here. Next.js does not read the repo-root.env, so the web app kept using sandbox while the workers were already on production — link tokens minted against one environment and exchanged against the other, with a confusing error. Sync, then restart both services.
Chase, Wells Fargo, Bank of America, Capital One and similar require Plaid to register your specific application with each bank before it can link. Submitted once from the dashboard, then reviewed per institution, and the reviews run in parallel. Budget 2–4 weeks for the large banks. Nothing in your code changes; the institution simply cannot be linked until its review clears. Smaller banks and credit unions generally work immediately.
The trial plan allows 10 live Items (one Item = one set of credentials at one institution). Unlinking does not return the slot. Link deliberately.
Two consequences worth knowing:
- A card that two people can both see is one card. Linking it under each
person's login burns two slots and double-counts every balance and
transaction — Plaid issues a different
account_idper Item, so the unique constraint never fires. The exchange endpoint refuses duplicates and names which household member already holds them. - Linking someone else's bank must be done as them: the link token is
scoped per household member, because Plaid keys its returning-user
experience off
client_user_id. Set Linking for before opening Link, or it opens straight into the owner's own Plaid session.
Access tokens are encrypted with APP_ENCRYPTION_KEY before they touch the
database, and are only ever decrypted inside the worker. The browser never
receives a Plaid token or secret. Account identifiers are masked to the last
four digits everywhere, including in anything sent to a model.
This repository is public and the app runs on real financial data, so the two must not meet.
-
scripts/hooks/commit-msgapplies the same checks to the commit message. The content check reads only the staged diff, so a commit whose code is clean can still publish real figures in its own prose — explaining why a heuristic was wrong is easiest with the numbers that broke it, and those numbers come off a statement. -
scripts/hooks/pre-commitblocks staged secrets (sk-ant-, JWTs, Google client secrets), email addresses, 9-or-more-digit numbers that look like account or reference numbers, account masks (‥1234,...1234), money written with cents (a round figure is usually a generic cap; cents come off a statement), and any term in.pii-denylist. Both hooks share the rules inscripts/hooks/pii-scan. -
Enable it on a fresh clone:
git config core.hooksPath scripts/hooks -
.pii-denylistis gitignored — the list of things you must not publish is itself something you must not publish. One term per line:echo "Some Landlord LLC" >> .pii-denylist -
Examples in comments are invented. Concrete examples explain a bug far better than abstract ones, so they stay concrete and stop being real.
-
Genuine false positive:
git commit --no-verify.
| Command | What it does |
|---|---|
npm run db:reset |
Re-applies all migrations to a clean DB (wipes data) |
npm run db:types |
Generates TS types from the live schema into packages/shared |
npm run env:sync |
Rewrites env files from supabase status |
npm run owner:create |
Creates/repairs the owner user (idempotent; OWNER_PASSWORD to rotate) |
npm run up |
Foreground boot, all platforms, no service manager |
All keys live in gitignored .env files, written by scripts/sync-env.mjs
and read server-side only. The browser sees only the anon key; RLS plus the
app_owner allow-list gate every row. Add future keys (Plaid, Anthropic,
Alpaca, Kalshi) to root .env — sync-env preserves unmanaged lines.
- Table grants: this Supabase version does not auto-grant DML on new
tables to API roles — every new table needs grants (see
supabase/migrations/20260804000300_grants.sql; default privileges now cover future tables). [auth.email] enable_signup = falsedisables email sign-IN entirely. Registration lockdown belongs to the global[auth] enable_signuponly.- Docker's daemon
"ip"option only affects the default bridge network. User-defined networks (compose, supabase CLI) needdefault-network-opts.bridge.com.docker.network.bridge.host_binding_ipv4, and the setting only applies to networks created after it — hence the full stop/start inscripts/harden-docker-loopback.sh. - The repo path contains a space, which pm2's daemon cannot survive
(unquoted shell paths). That's why process management is systemd user units
running
scripts/launch-*.jsinstead of pm2, and why those launchers exist rather than.binshims.
- Linking: Overview → "Link institution". Sandbox accepts any bank with
user_good/pass_good. The post-link dialog flags business accounts (auto-stamps every transaction + requests receipts). - Sync: every 6h automatically; "Sync now" on Overview queues a worker job. Recurring detection + net-worth snapshot run nightly at 02:00.
- Categorization: Plaid baseline → your rules (Transactions → Rules, with retroactive apply + preview) → merchant map (any inline category correction teaches it permanently) → review inbox for the leftovers.
- Self Improvement ingest: drop JSON/CSV/MD into
si-inbox/, orPOST /api/si/entrieswithAuthorization: Bearer $SI_API_TOKEN. - OAuth institutions need registration. Production approval alone does not
unlock the big banks. Plaid's Compliance Center → App profile (name, logo,
website URL, reason for data access, contact email) is submitted once, then
each OAuth institution reviews it independently — some take 2–4 weeks. Banks
that don't use OAuth work as soon as production keys are in
.env. Check which is which under View institutions in the dashboard. - Before linking real institutions (Trial plan, 10 Items, slots don't
return on unlink): purge sandbox data —
delete from institutions; delete from net_worth_snapshots; delete from recurring_items;via Studio or psql (cascades take accounts/transactions).
- Gmail receipts (multi-account). Account page → Connect Gmail, repeat
per mailbox. Each connection gets Choose receipt labels: chips listing
that inbox's real Gmail labels (via
/api/gmail/labels), plus a Purchases (auto) chip for Gmail's own purchase category. Selections build the poll query; a hand-written advanced query overrides them. Blank = the default receipt-subject search. A 7-day recency guard is auto-appended. Polling runs every 45s per mailbox; a duplicate guard (same vendor + total within ±6h) stops a CC'd receipt creating two anticipations. - Receipt parsing. Regex first (labeled totals, card last-four); when a
verified sender yields no total, an LLM pass classifies purchase-vs-noise
and extracts vendor/total/last4. Noise (shipping updates, mail previews) is
marked
ignoredand hidden. A background backfill re-parses older unparsed receipts a few per cycle. - Anticipation → reconciliation. Every legitimacy-verified receipt creates
an anticipated charge (card attributed via last-four ↔
accounts.mask). When the bank transaction posts, a composite score (amount, descriptor / name match, date, last-four) auto-reconciles at high confidence, raises a one-tap "Is X the same as Y?" prompt at medium (the answer persists as a permanent vendor alias), and leaves the rest in the ambiguous queue. Unposted after 14 days → review as possible refund / cancellation / fraud. - Watchlist. Flag any vendor
fraudorcancelled; hits fire at all four ingress points (receipt ingest, anticipation, Plaid sync, recurring detection) with an instant alert and an auto-filed dispute queue item. - Notifications. Set
NTFY_TOPICin.envand subscribe to that topic in the ntfy app for push alerts at email speed. - The agent. Daily 06:00, or Run analysis now on
/agent. Advisory until trading is switched on (see Turning on paper trading). Model viaAGENT_MODEL(defaultclaude-sonnet-5); receipt parsing viaRECEIPT_MODEL. - Budgets.
/budget— auto-fill from trailing 6-month averages, category and flex modes, rollovers, pace bars. - Reports.
/reports, five tabs sharing one period picker. Cash flow is a hand-built Sankey (income sources → cash in → spending groups → categories); the hub is deliberate, since nothing in the data says which paycheck paid which bill. Trends charts 12 months and compares the period against the prior period or the same period last year. Recaps reads the weekly/monthly recaps. Builder filters on anything and saves the configuration tosaved_reports. Business & tax exports business expenses by entity/category/date with a receipt reference on every line and a missing-receipt count. - Goals.
/goals— the wizard's three steps are define → link → preview. The linking step is the point: funding accounts, categories/tags that count as contributions, and cost-driver liabilities or recurring items. Step three replays six months of real transactions through that linkage and shows what last month's recap would have said, so a wrong link is visible before you save. Contributions auto-match nightly (both legs of a transfer collapse to one), manual attachments survive re-matching, and each goal's detail view expands every attributed cost down to its arithmetic. - Enrichment. Daily 05:30, before the agent runs. Transactions the
deterministic pipeline left uncategorized go to Claude in batches; ≥0.8
confidence applies the category, below that lands in the review inbox with
the model's own reason. Only single-purpose merchants at ≥0.92 teach
merchant_map— Amazon and Target never do, which is the whole point. The same pass scores personal-account transactions for business likelihood into a suggestion inbox on Transactions → Business; dismissing a merchant suppresses it permanently.ENRICH_MODELoverrides the model; Enrich now on/agentruns it on demand. - Recaps. Weekly Sunday 22:00, monthly on the 1st. Stage 1 is pure code: cash flow vs the prior period, budget adherence, credit utilization, goal cost attribution (interest from carried balance × APR ÷ 365 × days, stored with its formula and the exact contributing transaction ids), and net efficiency per goal. Stage 2 sends those facts to Claude for five 0–100 domain scores, the narrative, and adjustments. Every figure in the prose is then checked against Stage 1's numbers — an unsourced figure fails the whole run, which is logged and surfaced rather than shown. Accepting an adjustment files it into the Approval Queue; monthly recaps add subscription verdicts as keep/replace/cut/watch decision cards with an addressable-savings headline.
- Phase 0 — Foundation: local stack, schema, auth, worker skeleton, reboot persistence
- Phase 1 — Read everything: Plaid Link + sync (encrypted tokens), categorization pipeline v1, recurring detection, business layer v1, net-worth snapshots, SI section v1 — sandbox-verified; real-institution acceptance pending Trial approval
- Phase 2 — The brain (read-only): feature-complete except Aldyn (getaldyn.com) Path A
- Budgets (category + flex, rollovers, 6-month auto-fill)
- Agent worker v1 + Approval Queue (advisory) + Agent control page
- Gmail receipt ingestion (multi-mailbox, label mapping, LLM parse), anticipation engine, scored reconciliation, vendor watchlist
- Reports: cash flow Sankey, summaries, MoM/YoY, saved reports, business tax export
- Goals wizard +
goal_linkssemantic mapping + contribution matching + pace math - LLM categorization enrichment + business suggestion engine
- Recap engine (Stage 1 deterministic math → Stage 2 LLM scoring/narrative) + subscription review
- Aldyn (getaldyn.com) Receipts API (Path A) — replaces Gmail OAuth once live (blocked on the Aldyn-side build)
- Phase 3 — Human-approved execution: Alpaca paper → Kalshi demo → transfers, executor guardrails
- Executor + guardrails in code, append-only
executionsledger - Agent emits
tradeproposals, pre-checked against the same guardrails before they reach the queue - 30 days of paper trading with zero guardrail violations (needs broker keys)
- Investments page — positions with P&L and allocation, broker orders, and this platform's own attempts including refusals
- Kalshi demo (
prediction_position), then transfers
- Executor + guardrails in code, append-only
- Phase 4 — Bounded autonomy: allow-listed auto-execution, circuit breakers, notifications
- Phase 5 — Hosted migration (optional): Vercel + hosted Supabase + Render, same migrations
- Phase 5 — Desktop app packaging: Tauri shell, approval-badge count on the tray icon, native notifications wired to
notification_rules(tray v0 + launcher + installable PWA manifest exist now) - Settings: per-function model selector — choose the Claude model per LLM function (agent analysis, categorization enrichment, recap scoring, subscription review) from a settings UI; currently env-based (
AGENT_MODEL, defaultclaude-sonnet-5)
/chat answers questions about your own money — balances, the last 90 days of
spending, goals, recurring bills and your floors — from a snapshot rather than
from guesswork, with account identifiers masked to a last-four the same way the
agent's are. It is advisory: it cannot move money, place a trade, or change a
setting.
It runs on the Claude Code login already on the machine. No
ANTHROPIC_API_KEY, no metered billing. The Claude Agent SDK authenticates the
way Claude Code does, so a Pro or Max subscription is enough. What the honest
answer looks like per provider:
| Provider | Keyless? | Notes |
|---|---|---|
claude_code |
yes — uses the Claude Code login | chat only; no schema-enforced output |
anthropic |
no | ANTHROPIC_API_KEY; powers the workers |
google |
free tier, not a login | AI Studio issues a no-card key |
openai |
no | a ChatGPT subscription does not include API access; they are billed separately |
ollama |
yes — no account at all | local models, no network, no bill |
Chat and the workers are configured separately (app_settings.llm_chat_provider
versus llm_provider) and have to be. claude_code is chat-capable and
schema-incapable, and every worker depends on schema-enforced output — the
agent's recommendations, the recap's scores, categorisation, receipt parsing.
Pointing one shared setting at it would give you free chat and a nightly agent
that silently produces nothing, so resolveLlmSettings refuses to hand a
chat-only provider to a worker and falls back with a warning.
The chat can query the whole book, not just what fits in a prompt. It is
given a 90-day snapshot for orientation and three read-only tools for
everything else: list_accounts (every account, with its mask, balance, and for
credit its rate, statement balance and due date), search_transactions (any
date range, merchant, category, account or amount across the full history), and
spending_summary (totals grouped by category, merchant, month or account).
It can also recategorise transactions — the one thing it changes. It takes
transaction ids from a search rather than a filter, so what changes is exactly
what was looked at; it refuses an invented category name and any single
instruction touching more than 200 rows; and the previous categories go to the
audit log so a mistake can be undone rather than reconstructed. apply_to_future
teaches the merchant map, and merchants that sell across unrelated categories
are refused for that however explicitly asked — one Amazon refund filed by hand
once taught the map "Amazon = Refunds" and restated 537 purchases as income.
Handing it the history instead would be ~130k tokens per turn of mostly irrelevant context, and it still could not answer a question about one merchant in one month — whatever summary fits has already discarded the detail. To ask about one card, give its last four digits: several accounts share a name and differ only by the mask, so the digits are the only exact identifier.
Every filter runs in SQL. That sounds like an implementation note and is not: an
earlier version applied limit in the database and filtered by account
afterwards, so asking for one card returned the most recent rows across all
accounts and kept whichever happened to match — reporting "0 transactions" for a
card with a four-figure statement balance, confidently. A wrong answer delivered
calmly is worse than an error when a model is reasoning from it.
The tools run under the owner's own session, so RLS applies to them exactly as it does to the pages.
Semantic merchant recall (find_merchants) answers what a substring cannot:
"that garden place", "the moving company", "streaming subscriptions". Embeddings
run locally through Ollama (nomic-embed-text, 768 dimensions) into pgvector,
so a list of everywhere the owner shops never leaves the machine — which for
this data is not a nice-to-have. scratch/build-merchant-index.mts builds it;
only merchants whose text changed are re-embedded, so a rebuild is ~2s once
steady rather than ~19s cold.
Merchants are indexed, not transactions: 785 against 3,099, and every transaction from one merchant embeds to the same point anyway. And the source text is the name plus its dominant category, not the name alone — a bare "1-800-Pack-Rat" contains no word about moving, so searching "moving and storage company" ranked a wine shop with "Warehouse" in its name above it. With the category attached, "electric utility" went from 0.70/0.58/0.55 to 0.79/0.79/0.78 across three real utilities, and "coffee shop" stopped returning a smoke shop.
This is a recall feature, not a cost saving. The token problem was solved by querying rather than carrying history; embeddings do not shrink what is left. Its remaining misses are mostly data quality rather than retrieval — a moving company filed under Other Income cannot be found by describing what it does.
The chat gets 30 steps per turn (CHAT_MAX_TURNS). Categorising is list,
search, write, re-check, summarise, then answer, and an earlier budget of eight
ran out before reaching the answer — discarding the whole turn while any
categorising it had already done stayed committed and unreported. Running out
now returns whatever was written plus a plain account of where it stopped,
because a loop that can write must never fail silently mid-way.
Measured, not assumed: the per-turn snapshot is ~1,650 tokens. Carrying the
full history instead would be ~132,000. The snapshot was 5,538 until rates,
statement balances and the full recurring list moved into list_accounts — 59%
of it was interest-rate detail answering a question asked on maybe one turn in
ten. scratch/measure-chat-tokens.mts prices each section if it drifts again.
Model and effort are set per surface from the gear on /chat. The model
list is asked of the SDK rather than hardcoded, so it reflects whatever Claude
Code ships with — and the effort choices are the ones that model actually
accepts. That matters because the rules are not uniform: effort is rejected
by Haiku 4.5 rather than ignored, budget_tokens returns a 400 on Sonnet 5 and
Opus 5/4.8/4.7 while Opus 4.6 still takes it, and Opus 5 only permits thinking
to be switched off at effort high or below. A level a model does not support
is stepped down or omitted before the request is built, so changing the model
never turns a saved setting into an error.
Images and PDFs can be attached — dropped, pasted, or picked. Bytes go to
the private chat-attachments bucket (already covered by npm run db:backup)
and the metadata to Postgres, so reopening a thread months later still shows
what was asked about. Images or PDF only, 20 MB each; unsupported files are
refused with a reason rather than dropped silently. Where a provider cannot read
a PDF — OpenAI's chat endpoint, local Ollama models — the model is told a PDF
was attached rather than left to answer as if nothing were there.
Interest rates and statement balances are on the Overview beside each credit account, and in the chat's context. A card carries several rates at once — purchases, cash advances, balance transfers, and a promotional rate that overrides the others while it lasts — so the rate being paid is shown with the rest on hover. Plaid does not report rates for every issuer (9 of 16 here), and those read "rate not reported" rather than blank, because a missing rate should never look like 0%.
The statement balance is shown as its own figure, distinct from the balance above it. They are different numbers: the statement balance is what must be paid to avoid interest, while the current balance includes everything charged since the statement closed and is not yet owed. Conflating them is how this platform once reported interest accruing on a card that is paid in full every month.
The Agent SDK is Claude Code, which means it ships with a filesystem and a
shell. Every tool is denied, the Claude Code system preset is replaced, and
settingSources: [] keeps your CLAUDE.md and permission rules out of a context
that has no business seeing them. What is left is a model call with a
subscription behind it.
By default nothing on the network can reach this app — it binds loopback only.
WEB_HOST in the root .env changes that. Prefer binding the VPN address
specifically (WEB_HOST=100.x.y.z) over 0.0.0.0: it serves the tunnel and
nothing else, so the app is never available over plain HTTP on the local
network even for a moment.
The session cookie name is pinned (sb-life-command-auth-token). Supabase
otherwise derives it from the Supabase URL — sb-<first host label>-auth-token
— which is invisible until the browser and the server disagree about that URL.
Once the browser talks to 100.x.y.z:3141 and the server to 127.0.0.1:54321,
sign-in succeeds and issues a real token, then the next request carries a cookie
the server never looks for and bounces back to the login page. Nothing errors;
it reads exactly like a wrong password. Pinning it also makes a session portable
across localhost, LAN and VPN addresses.
Supabase is not opened alongside it. The browser reaches the database
through this app at /supabase, proxied to a Supabase that stays bound to
localhost. That is not just tidiness: NEXT_PUBLIC_SUPABASE_URL is compiled
into the browser bundle as 127.0.0.1:54321, and on a phone that address means
the phone. The page would load and every login and query would fail against a
database that isn't there. The browser client builds its URL from whatever
origin served the page, so one build works from localhost, a LAN address, or a
Tailscale name with nothing to reconfigure.
The proxy path is excluded from the auth middleware. Left in, the session check
treats every auth call as an unauthenticated page request and returns a login
page to a fetch — which presents as a password that simply doesn't work.
WEB_HOST=0.0.0.0 serves plain HTTP. Session cookies and the six-digit code
you type at sign-in cross the network unencrypted, readable by anyone who has
your Wi-Fi password — which on a home network includes guests and every smart
device on it. For a page showing bank balances and holding encrypted Plaid
tokens, that is a poor trade for convenience.
A mesh VPN fixes it without exposing anything publicly: install
Tailscale on this machine and on the phone,
sign both into the same account, and reach the app at
http://<machine-name>:3141 over an encrypted WireGuard link. It works away
from home as well, needs no port forwarding, and nothing is published to the
internet. Keep WEB_HOST=0.0.0.0 — Tailscale presents its own interface, so the
app must listen on more than loopback either way.
Do not put this behind a public tunnel. A URL anyone can reach is a URL anyone can attack, and the only thing between it and your accounts would be one password and one TOTP.
The database is the only thing here that cannot be rebuilt. Transactions re-sync from Plaid and everything derived from them recomputes; what does not come back is the teaching — manual corrections and the merchant map built from them, review decisions, goals, budgets, household assignments, receipts, and the Plaid access tokens, which live encrypted in the database and cannot be re-issued by Plaid. Losing them means re-linking every institution, and on a Trial plan those Item slots are not returned.
This install learned that on 2026-08-29. The repo was being synced off-machine the whole time; Postgres was not, because it lived in a Docker volume on a scratch disk outside every synced path.
npm run db:backup # pg_dump -Fc + storage bucket, dated
npm run db:restore -- --list # what is available
npm run db:restore # newest, after confirmation
npm run db:restore -- <file> # a specific dump
Backups are written outside the repo, to a sibling directory
(../Supabase Backup - Finance Dashboard/, override with BACKUP_DIR). That is
deliberate: a dump holds real balances and merchant names and this repo is
public, so keeping it out of the working tree means it cannot be committed even
by accident. Whatever already syncs your projects directory carries it
off-machine without further configuration.
Each run writes three files plus the bucket when it has contents — the dump,
a pg_restore -l table of contents, and a log recording row counts at the time
of the dump. The row counts matter: a dump that restores cleanly and contains
nothing is the failure this exists to prevent, and a file size in megabytes does
not tell the two apart. Fourteen sets are kept (BACKUP_KEEP).
npm run svc:install installs finance-backup.timer, which runs nightly at
03:30 with Persistent=true so a machine that was asleep still gets its backup
when it wakes.
Restore is written and tested alongside the backup, not after the first
disaster. A dump nobody has restored is a file, not a backup. It was verified
the only way that counts — drop schema public cascade, then restore — which is
what caught the dump omitting CREATE SCHEMA, and --no-acl silently stripping
every GRANT.
A failed backup fires finance-backup-notify.service: a desktop toast and, if
NTFY_TOPIC is set, a push. A nightly job that stops running looks exactly like
one that runs, and the failure that costs you something is the one nobody was
sitting at the machine to see.
What this still does not protect against, so it is not mistaken for more than it is:
- Up to 24 hours of loss between snapshots. Transactions re-sync from Plaid, so the real exposure is a day of manual corrections.
- The backups share a disk with the repo (
/), not with the database (/mnt/scratch) — which is what makes them survive the failure that has already happened once. A root-disk failure is a different story, and the off-machine copy in Drive is what covers it. - Nothing yet proves an old dump still restores. The nightly run checks that a dump was written and records row counts; it does not restore it. A periodic test-restore into a throwaway database is the honest next step.
Aldyn (getaldyn.com) is a separate receipt-capture product by the same author. This platform can take receipts from it directly (Path A) instead of scraping Gmail (Path B), which removes the OAuth dependency and gets structured line items rather than parsed HTML. Path A is optional and not required to run anything here — Gmail ingestion works standalone, and everything in this repository is written against Path B.
Everything lives in the gitignored root .env; npm run env:sync propagates
what the web app and workers need.
| Key | Needed for | Status |
|---|---|---|
SUPABASE_*, DATABASE_URL |
local stack | managed by env:sync |
APP_ENCRYPTION_KEY |
Plaid/Gmail token encryption | generated at setup |
SI_API_TOKEN |
POST /api/si/entries |
generated at setup |
PLAID_CLIENT_ID / PLAID_SECRET / PLAID_ENV |
banking aggregation | set (production) |
ANTHROPIC_API_KEY |
agent + receipt parsing | set |
GOOGLE_CLIENT_ID / GOOGLE_CLIENT_SECRET |
Gmail receipts | set |
AGENT_MODEL / RECEIPT_MODEL / ENRICH_MODEL / RECAP_MODEL |
per-function model overrides (default claude-sonnet-5) |
optional |
NTFY_TOPIC |
push notifications (recaps, receipts, watchlist) | optional |
WEB_HOST |
bind address for the web app — 0.0.0.0 to reach it from other devices |
defaults to loopback |
ALPACA_KEY_ID / ALPACA_SECRET_KEY |
Phase 3a paper trading | needed to execute |
ALPACA_PAPER_BASE |
points paper at a stub broker instead of Alpaca, for exercising order placement without an account | testing only |
KALSHI_* |
Phase 3b | not yet needed |
NTFY_TOPIC is a credential, not a name: on ntfy.sh anyone who knows the topic
string reads every notification it carries. Generate a random one and treat it
like a password.
NTFY_TOPIC=lifecmd-$(openssl rand -hex 16)
A fresh install refuses every order, and that is the intended starting state.
Trading is opt-in at five independent points, each checked against the world
rather than asserted — turning on four of them does nothing. The Execution
card on /agent lists all five and says which one is off, since an empty
approval queue looks identical either way.
ALPACA_KEY_IDandALPACA_SECRET_KEYin root.env, thennpm run env:sync. This one is the only step that is not in the UI — secrets stay in the environment.- Allow-list
trade(agent_config.allowed_action_types). - Autonomy level —
1lets the agent propose trades,2lets an approved one execute.0is read-only and disarms both. - The caps, all defaulting to
0: per-transaction, daily, per-position, and open positions. A cap of zero refuses everything. - One agent-controlled account. Alpaca is not a bank Plaid aggregates, so Create it from the broker makes the row by asking Alpaca who it is. Zero flagged accounts means no trading; two or more also means no trading, because the agent must never pick between accounts — code stamps the account into every order and the model never chooses. Naming one unnames the rest.
agent_config.execution_mode defaults to paper and must be changed by hand;
the spec gates live on 30 days of clean paper trading.
Both sides check the same rules. The agent discards a proposal that would breach
a cap before the owner ever sees it, so the queue only offers Approve on
something that could actually go through; the executor then re-derives every
input from the database and the broker and decides again at execution time,
because approval says "I want this" and says nothing about whether it is still
inside the limits. Every attempt is written to executions, refusals included —
"30 days with zero guardrail violations" is not checkable against a table that
only remembers the orders that went through.
Google OAuth setup: Cloud Console → enable Gmail API → OAuth consent screen
(External, published unverified so refresh tokens don't expire weekly) →
credentials → Web application → redirect URI
http://localhost:3141/api/gmail/callback.