diff --git a/openspec/changes/csv-backfill-import/.openspec.yaml b/openspec/changes/archive/2026-07-16-csv-backfill-import/.openspec.yaml similarity index 100% rename from openspec/changes/csv-backfill-import/.openspec.yaml rename to openspec/changes/archive/2026-07-16-csv-backfill-import/.openspec.yaml diff --git a/openspec/changes/csv-backfill-import/design.md b/openspec/changes/archive/2026-07-16-csv-backfill-import/design.md similarity index 100% rename from openspec/changes/csv-backfill-import/design.md rename to openspec/changes/archive/2026-07-16-csv-backfill-import/design.md diff --git a/openspec/changes/csv-backfill-import/proposal.md b/openspec/changes/archive/2026-07-16-csv-backfill-import/proposal.md similarity index 100% rename from openspec/changes/csv-backfill-import/proposal.md rename to openspec/changes/archive/2026-07-16-csv-backfill-import/proposal.md diff --git a/openspec/changes/csv-backfill-import/specs/csv-import/spec.md b/openspec/changes/archive/2026-07-16-csv-backfill-import/specs/csv-import/spec.md similarity index 100% rename from openspec/changes/csv-backfill-import/specs/csv-import/spec.md rename to openspec/changes/archive/2026-07-16-csv-backfill-import/specs/csv-import/spec.md diff --git a/openspec/changes/csv-backfill-import/specs/reporting/spec.md b/openspec/changes/archive/2026-07-16-csv-backfill-import/specs/reporting/spec.md similarity index 100% rename from openspec/changes/csv-backfill-import/specs/reporting/spec.md rename to openspec/changes/archive/2026-07-16-csv-backfill-import/specs/reporting/spec.md diff --git a/openspec/changes/csv-backfill-import/tasks.md b/openspec/changes/archive/2026-07-16-csv-backfill-import/tasks.md similarity index 100% rename from openspec/changes/csv-backfill-import/tasks.md rename to openspec/changes/archive/2026-07-16-csv-backfill-import/tasks.md diff --git a/openspec/specs/csv-import/spec.md b/openspec/specs/csv-import/spec.md new file mode 100644 index 0000000..59e5d10 --- /dev/null +++ b/openspec/specs/csv-import/spec.md @@ -0,0 +1,255 @@ +# csv-import Specification + +## Purpose +TBD - created by archiving change csv-backfill-import. Update Purpose after archive. +## Requirements +### Requirement: Import attribution to an existing account + +The system SHALL require every CSV import to be attributed to exactly one account +that already exists in the database. The system SHALL NOT create, rename, or +otherwise modify an account as a result of an import, preserving the +`account-management` rule that accounts originate only from sync data. Accounts in +state `HIDDEN` SHALL NOT be offered as import targets. + +#### Scenario: User selects a target account + +- **WHEN** a user begins a CSV import +- **THEN** the system requires them to choose one existing non-hidden account, and + every transaction ingested from that file is attached to that account + +#### Scenario: No account selected + +- **WHEN** a user attempts to proceed without choosing a target account +- **THEN** the import does not proceed and no rows are ingested + +### Requirement: Verbatim archival before normalization + +The system SHALL store the uploaded file's bytes verbatim in an `imports` record +before any parsing, mapping, or normalization occurs. The archived payload SHALL +never be mutated or deleted by the application. The system SHALL store the chosen +column mapping and the duplicate decisions on the same record, so that the +outcome of an import is a deterministic function of the archived record alone. + +#### Scenario: File archived on upload + +- **WHEN** a user uploads a CSV file +- **THEN** an `imports` row is committed containing the exact uploaded bytes with + status `draft`, before any row is parsed + +#### Scenario: Import is replayable + +- **WHEN** the import logic is re-run over an archived record with its stored + mapping and decisions +- **THEN** it produces the same set of transactions as the original run + +#### Scenario: Unparseable file + +- **WHEN** an uploaded file cannot be parsed as CSV +- **THEN** the archived record is retained, the failure is stated plainly with the + reason, and no transactions are ingested + +### Requirement: Column mapping + +The system SHALL detect a CSV's date, amount, and description columns from its +header row where possible, and SHALL require the user to confirm or correct the +mapping before commit. The system SHALL support amounts expressed as a single +signed column and as separate debit and credit columns. The system SHALL treat the +interpretation of ambiguous numeric dates (whether `03/04/2026` is March 4th or +April 3rd) as a user-confirmable part of the mapping, inferring it from the data +where the data settles it and defaulting to month-first otherwise. The system +SHALL accept the negative conventions common to bank exports, including +parenthesized values and trailing minus signs, and SHALL convert amounts to +integer cents without floating-point arithmetic. The system SHALL map the file's date to the +transaction's posted timestamp and record imported transactions as not pending; +when the file exposes a distinct transaction date, the system SHALL map it to the +transaction date. + +#### Scenario: Headers auto-detected + +- **WHEN** a CSV's header row contains recognizable date, amount, and description + columns +- **THEN** the system pre-selects them and presents the mapping for confirmation + +#### Scenario: User corrects a wrong guess + +- **WHEN** the auto-detected mapping is wrong +- **THEN** the user can reassign any field to any column before committing + +#### Scenario: Date order inferred from the data + +- **WHEN** a file's date column contains `13/04/2026`, which can only be + day-first +- **THEN** the whole file is interpreted day-first + +#### Scenario: Ambiguous date order is confirmable + +- **WHEN** every date in a file is ambiguous, such as `03/04/2026` +- **THEN** the system defaults to month-first, states the resulting date range + before commit, and allows the user to select day-first instead + +#### Scenario: Parenthesized negative + +- **WHEN** an amount column contains `(12.34)` +- **THEN** it is ingested as `-1234` integer cents + +#### Scenario: Separate debit and credit columns + +- **WHEN** a file expresses amounts as separate debit and credit columns +- **THEN** the user can map both, and each row resolves to a single signed integer + cent amount + +#### Scenario: Imported rows appear in reports + +- **WHEN** an imported transaction is committed +- **THEN** it is recorded as not pending with its posted timestamp set, and + appears in monthly reports for the month of that timestamp + +### Requirement: Stable synthetic identity + +The system SHALL derive a stable, deterministic identifier for each imported +transaction from its content — the target account, date, amount, and description — +combined with an occurrence index distinguishing rows that are otherwise +identical within the same file. Identifiers SHALL be namespaced so they are +distinguishable from provider-supplied identifiers. Importing a file whose rows +have already been ingested under the same identifiers SHALL NOT create duplicate +rows. + +#### Scenario: Same file imported twice + +- **WHEN** a user imports a file and then imports the identical file again +- **THEN** no new transactions are created and the second import reports that + every row already exists + +#### Scenario: Overlapping files + +- **WHEN** a user imports a January–March file and then a February–April file into + the same account +- **THEN** the February–March rows are recognized as already ingested and only the + April rows are added + +#### Scenario: Genuinely identical transactions + +- **WHEN** a file contains two rows with the same date, amount, and description, + representing two real transactions +- **THEN** both are ingested as separate transactions + +### Requirement: Potential duplicate review + +The system SHALL, before commit, identify each candidate row that would create a +transaction resembling one the target account already holds — matching on +identical amount and an effective date within one day, against non-removed +transactions regardless of their origin. The system SHALL NOT match on +description, because the same transaction is rendered differently by different +sources. The system SHALL present each flagged pair to the user with both +descriptions shown, and SHALL default the flagged row to being skipped. The user +SHALL be able to override any flagged row to be imported. + +#### Scenario: Cross-source duplicate flagged + +- **WHEN** a candidate row has the same amount and date as an existing synced + transaction in the target account, but a differently worded description +- **THEN** the row is flagged for review, both descriptions are shown side by + side, and it defaults to skipped + +#### Scenario: Default skip retains the synced row + +- **WHEN** a user commits an import without changing any duplicate decision +- **THEN** every flagged row is skipped, the existing transactions are left + untouched, and no duplicate is created + +#### Scenario: User keeps a false positive + +- **WHEN** a flagged row is in fact a distinct transaction and the user marks it to + be imported +- **THEN** it is ingested as a new transaction alongside the existing one + +#### Scenario: Preview before commit + +- **WHEN** a mapped file is ready for review +- **THEN** the system states the parsed date range, the number of rows to be + imported, the number already present, and the number flagged as potential + duplicates + +### Requirement: Additive-only ingestion + +An import SHALL only insert transactions. The system SHALL NOT, as a result of an +import, modify or remove any existing transaction, reconcile pending transactions, +change any account's state, or write balance snapshots. Reconciliation and +feed-authority behavior SHALL remain exclusive to sync. + +#### Scenario: Pending transactions untouched + +- **WHEN** an import is committed into an account that holds pending transactions + absent from the CSV +- **THEN** those pending transactions remain unchanged and are not removed + +#### Scenario: Existing transaction not overwritten + +- **WHEN** a candidate row resolves to an identifier already present in the account +- **THEN** the existing transaction is left exactly as it was + +#### Scenario: No balance snapshots + +- **WHEN** an import is committed +- **THEN** no balance snapshot is written and the net worth report is unaffected + +#### Scenario: Account state unaffected + +- **WHEN** an import is committed into an `ACTIVE` account +- **THEN** the account's state and `last_successful_data_at` are unchanged + +### Requirement: Rule application after import + +The system SHALL apply existing categorization rules to imported transactions once +the import is committed, using the same rule precedence and the same `rule` +categorization event source as sync. Imported transactions SHALL NOT introduce a +new categorization event source. + +#### Scenario: Backfilled history categorizes itself + +- **WHEN** an import is committed and existing rules match some of the imported + transactions +- **THEN** those transactions are categorized with `rule` events recording the + winning rule, exactly as if they had arrived by sync + +#### Scenario: Manual decisions respected + +- **WHEN** rules are applied after an import +- **THEN** transactions whose latest categorization event is `manual` are not + re-categorized + +### Requirement: Import record and undo + +The system SHALL record every import and surface the record in Settings, showing +the file name, target account, parsed date range, counts of rows imported and +skipped, and the commit time. Each imported transaction SHALL record which import +produced it. The system SHALL allow a committed import to be undone as a unit, +removing the transactions it produced without deleting any categorization event +history. An undone import SHALL be re-importable after correcting its mapping. + +#### Scenario: Import listed in settings + +- **WHEN** a user opens the import record in Settings +- **THEN** each import is listed with its file name, account, date range, counts, + and commit time + +#### Scenario: Undo removes only that import's rows + +- **WHEN** a user undoes an import +- **THEN** exactly the transactions that import produced are removed from the + ledger and reports, no synced transaction is affected, and the import is marked + undone + +#### Scenario: Event history survives undo + +- **WHEN** an imported transaction had been categorized and its import is undone +- **THEN** the transaction no longer appears in the ledger or reports, and its + categorization events remain in the append-only log + +#### Scenario: Re-import after a corrected mapping + +- **WHEN** a user undoes an import made with a wrong mapping and re-imports the + same file with a corrected mapping +- **THEN** the corrected transactions are ingested and the undone import's rows do + not reappear + diff --git a/openspec/specs/reporting/spec.md b/openspec/specs/reporting/spec.md index a84cc0a..a1601f6 100644 --- a/openspec/specs/reporting/spec.md +++ b/openspec/specs/reporting/spec.md @@ -30,7 +30,7 @@ The system SHALL display net worth over time computed from balance snapshots: fo - **THEN** its balances are excluded from the net worth series ### Requirement: Transaction ledger -The system SHALL provide a ledger view of transactions filterable by account, category (including uncategorized), month, and pending status, and searchable by free text. Text search SHALL match case-insensitively against the raw description, payee, and memo fields and against the effective displayed name produced by rule display-name overlays, so search finds what the user sees. Search SHALL compose with all structured filters. The ledger shows date, account, displayed name (rule overlay applied when present), amount, category, and a provenance indicator (rule vs. person). The ledger is the surface for manual categorization and for creating rules from transactions. +The system SHALL provide a ledger view of transactions filterable by account, category (including uncategorized), month, pending status, and source (synced vs. imported), and searchable by free text. Text search SHALL match case-insensitively against the raw description, payee, and memo fields and against the effective displayed name produced by rule display-name overlays, so search finds what the user sees. Search SHALL compose with all structured filters, and the source filter SHALL compose with all other filters. The ledger shows date, account, displayed name (rule overlay applied when present), amount, category, and a provenance indicator (rule vs. person). The ledger SHALL NOT mark a transaction's source on the row itself; a transaction's origin — synced from a connection, or imported from a named file — SHALL be shown in its history panel. The ledger is the surface for manual categorization and for creating rules from transactions. #### Scenario: Filter to uncategorized - **WHEN** a user filters the ledger to uncategorized transactions @@ -48,6 +48,18 @@ The system SHALL provide a ledger view of transactions filterable by account, ca - **WHEN** a user searches for "netflix" with a month filter active - **THEN** only that month's transactions matching the text (raw fields or displayed name) are listed +#### Scenario: Filter to imported transactions +- **WHEN** a user filters the ledger by source to imported transactions +- **THEN** only transactions produced by a CSV import are listed + +#### Scenario: Source filter composes +- **WHEN** a user filters by source to imported with an account and month filter active +- **THEN** only that account's imported transactions in that month are listed + +#### Scenario: Origin shown in history +- **WHEN** a user opens the history panel of a transaction that came from a CSV import +- **THEN** the panel states that it was imported and names the file and import date, distinguishing it from a transaction synced from a connection + ### Requirement: Dashboard month-to-date overview The system SHALL display on the dashboard, scoped to the current calendar month and to non-hidden accounts: (1) a spending pie chart of expense totals by category computed from posted transactions with the same semantics as the monthly report (transfer-kind categories excluded, uncategorized shown as its own slice); (2) total income and total expenses from posted transactions; (3) the count and total amount of pending transactions whose effective date falls in the current month; and (4) a small fixed number of the most recent transactions, each linking to the transaction ledger. Each section SHALL render a clear empty state when the month has no qualifying data.