179 lines
11 KiB
Markdown
179 lines
11 KiB
Markdown
# Stock & Portfolio Analyzer
|
||
|
||
Interactive tool for analyzing individual securities and portfolios
|
||
against local Yahoo Finance dumps (`~/prog/fin/stocks`, ~4k symbols).
|
||
|
||
## Quick start
|
||
|
||
```bash
|
||
python3 -m venv .venv
|
||
.venv/bin/pip install -r requirements.txt
|
||
./run.sh # serves the UI on the fixed port 8599 (http://localhost:8599)
|
||
```
|
||
|
||
First run builds a parquet cache in `.cache/` (~1 min for 4k symbols);
|
||
later runs load in well under a second. The cache tracks the data dir
|
||
per-file (mtime + size in `.cache/manifest.json`), so when the download
|
||
is updated, only the changed/added/removed symbols are re-read — a
|
||
partial refresh takes seconds instead of a full ~1 min rebuild.
|
||
|
||
## Modules
|
||
|
||
| Module | Purpose |
|
||
|-----------------|---------|
|
||
| `data.py` | Ingest `{sym}-history/dividend/capitalGain.csv` -> cached parquet panels (date x symbol), with manifest-based incremental refresh when the data dir changes. `Adj Close` already includes distributions, so it drives pre-tax total returns. |
|
||
| `metrics.py` | Total/annualized return, vol, Sharpe, Sortino, max drawdown, Calmar, CAPM beta/alpha. Pure pandas, all transparent. |
|
||
| `portfolio.py` | Weighted portfolios with drift and periodic rebalancing to target weights (`1W/1ME/QE/YE`), one-way cost in bps. Spec grammar: commas join the elements of ONE portfolio (`SYM` or `SYM:w`, bare = equal weight), spaces separate DISTINCT symbols/portfolios (`parse_items`). |
|
||
| `tax.py` | Simplified DAS after-tax engine: FIFO lots, 365-day long/short split, separate LT/ST/dividend rates plus federal NIIT and a state+local marginal rate (applied at ordinary rates — state/local have no preferential cap-gain rate). Headline curve = what you keep if you **sell everything today** (unrealized gains taxed daily by lot age). |
|
||
| `chart_widget.py` | Self-contained plotly.js chart in an iframe: mouse zoom/pan, x clamped to the data, view edges snapped to first/last data points with day-precise labels, y tight-fit, every line re-based to 1.0 at the left edge. |
|
||
| `portfolios.py` | Saved portfolio definitions in `portfolios.json` (name, spec, scheme, cost). |
|
||
| `settings.json` | Persisted UI inputs (symbol/benchmark specs, scheme, costs, tax rates, period, curve/window mode) — restored on every page load and server restart; delete to reset. |
|
||
| `app.py` | Streamlit UI: single "symbol or portfolio" spec field (page updates as soon as the input is valid; unknown symbols get click-to-fix "did you mean" suggestions) + a benchmark box with the same grammar (one benchmark per line; a line is a single symbol or a comma-joined portfolio, simulated with the same scheme/cost/tax rules — pre- and after-tax curves, first one drives beta/alpha), scheme/costs/tax rates, save + load/compare/delete portfolios (overlaid pre/after-tax curves), curve toggle (both / pre-tax only / after-tax only), stats table, allocation, per-year tax detail. |
|
||
|
||
## Data verification (Tier 3: against official filings)
|
||
|
||
Notes on the data itself:
|
||
|
||
- **Capital gains in this dump are mutual funds only.** The
|
||
dividend/capitalGain event files exist only for open-end mutual funds
|
||
(spot-checked: no ETFs, CEFs, BDCs or individual stocks carry them —
|
||
CEF/BDC distribution components are lumped into their dividend stream
|
||
by Yahoo). For total returns only the distribution TOTAL matters, which
|
||
is what Yahoo provides and what the corrections above fix.
|
||
- **There is no long-term/short-term split anywhere in the Yahoo data**
|
||
(an event is just a per-share amount), and the corrections never used
|
||
one either — the official checks compare total distributions. `tax.py`
|
||
therefore taxes capital-gain distributions at the long-term rate
|
||
(`lt_rate`), a simplification: fund cap-gain distributions are
|
||
predominantly long-term, but the per-fund LT/ST split would come from
|
||
the fund company's annual tax statement (1099-DIV detail: boxes 2a/2b),
|
||
the shareholder-report body, or a commercial feed (Lipper/Morningstar).
|
||
- **NYC residents: add the state+local layer.** The sidebar's two extra
|
||
rate fields exist for this. Set `NIIT %` to 3.8 if your MAGI is over
|
||
the threshold, and `State + local %` to your NY+NYC MARGINAL rate sum
|
||
(2025 single, from the Form IT-201 rate schedules): NYC is 3.876%
|
||
above $50k of city taxable income; NY is 6.85% at $215,400-$1.077M of
|
||
state taxable income and 9.65% at $1.077M-$5M. So a typical NYC
|
||
household pays, on a capital-gain distribution, roughly
|
||
20% federal + 3.8% NIIT + 6.85-9.65% NY + 3.876% NYC = ~34-37% —
|
||
which is exactly why high-distribution open-end funds lose so much
|
||
more to tax than ETFs for NYC residents (see the per-year tax detail
|
||
tab). Note: NY/NYC tax capital gains at ORDINARY rates (no
|
||
preferential cap-gain rate), which the single `State + local %` field
|
||
models correctly; the federal `Long-term gains %` field stays the
|
||
preferential 0/15/20%.
|
||
|
||
`~/prog/fin/stocks` is a Yahoo dump and gets re-downloaded (overwritten),
|
||
so fixes must live outside it. Pipeline:
|
||
|
||
1. `scripts/audit_stale.py` — finds tickers whose latest history row is
|
||
old and whose fresh Yahoo download is empty (delisted/merged funds,
|
||
tickers Yahoo no longer serves). Snapshots their final
|
||
`{history,dividend,capitalGain}.csv` + `longName` into
|
||
`overrides/frozen/{SYM}.{ext|json}`. `data.py` shadows the data root
|
||
with the frozen copies, so a re-download can't clobber them.
|
||
1b. `overrides/event-backup/` — backup of every dividend/capitalGain file
|
||
(refresh with `scripts/backup_events.py` after each dump update).
|
||
Yahoo changed its event feed in 2026: for some funds it no longer
|
||
returns `capitalGains` events at all, and terminated funds return no
|
||
events, so a re-download can empty or orphan good event history. The
|
||
data-root file wins while it is populated; an emptied/missing one
|
||
falls back to the backup (`data.py event_file`). The `ohlc` converter
|
||
likewise refuses to overwrite a populated event CSV with an empty
|
||
download. (Yahoo omits the whole `capitalGains` JSON key when there
|
||
are no events, so plain re-downloads in place usually leave old files
|
||
untouched — the backup is the second line of defense.)
|
||
2. `scripts/verify_official.py [syms | --stale]` — finds the fund's
|
||
shareholder report (EDGAR EFTS for `"TICKER"`, forms N-CSR/N-CSRS/
|
||
N-14/N-2/497) and parses the per-class *Financial Highlights* (or
|
||
JPMorgan-style *Per share operating performance*) tables: period-by-
|
||
period distribution totals compared against local
|
||
dividends + capital-gains over the same windows, plus a spot check of
|
||
the NAV-per-share row against the local close. Verdicts per symbol in
|
||
`reports/xcheck_official/{sym}.json`: **ok** (all bounded fiscal
|
||
years agree), **mismatch** (class matched, some year off — the
|
||
report tells you which), **weak-match** (best class fit too poor to
|
||
trust), **not-found** (fund not in any candidate filing). The local
|
||
side applies the corrections overlay, so a corrected fund verifies
|
||
against its filing. When several filings parse (e.g. the 497 annual
|
||
and the N-CSRS, which can round differently), the best match wins.
|
||
3. Confirmed findings go into `overrides/corrections/{SYM}.json` as
|
||
auditable deltas (`remove`/`replace`/`add` of distribution rows, each
|
||
entry dated and valued — amounts are what Yahoo reports, not the
|
||
official filing's). `data.py` applies them on cache build. Examples:
|
||
- **CVSIX** — 2008-12-18 0.292 duplicate of the same day's 0.641
|
||
(official FY09 = 0.81 balances without it); 2023-12-21 0.510
|
||
capital-gain row duplicated next to the day's 0.691 dividend
|
||
(official FY2024-10 = 0.79 balances without it).
|
||
- **JLPSX** — 13 year-end capital-gain distributions duplicated into
|
||
the dividend file (FY2021–2025 all match the JPMorgan 497 after
|
||
removal).
|
||
|
||
Remaining known data gaps (need the fund company's per-date
|
||
distribution archive; fiscal-year totals alone can't reconstruct them):
|
||
FAEVX and FGIZX — the Yahoo dump is missing the funds' regular quarterly
|
||
dividend rows (official fiscal-year totals exceed local even after the
|
||
double-listing dedup above); FIKAX — the official extraction is ambiguous
|
||
(the 500-fund consolidated Fidelity N-CSRS has multiple near-identical
|
||
sub-fund tables, and the best match is a systematic ~0.12/yr offset,
|
||
i.e. the wrong share class).
|
||
|
||
**Whole-dump double-listing sweep.** `scripts/scan_double_listing.py`
|
||
scans every symbol's effective event files for the pattern and writes
|
||
`reports/double_listing/scan.{md,json}`. On the 2026-08 dump (6,421
|
||
symbols with event files): 1,218 symbols had same-date/same-amount
|
||
cross-file pairs (6,589 pairs, 80% in December / fiscal year-end), 1,631
|
||
had same-date differing-amount pairs (12,366), and 83 had repeated
|
||
same-date rows within one file (Yahoo repeating a row up to ~35x; the
|
||
ingest already keeps the last). The same-amount pairs were corrected in
|
||
bulk by `scripts/apply_dedup_corrections.py` (1,212 correction files,
|
||
`dedup` invariant, mechanism-inferred and individually revertible). The
|
||
differing-amount pairs are reported but NOT auto-corrected: without a
|
||
per-fund official schedule they can't be distinguished from a legitimate
|
||
same-day dividend + capital-gain pairing — that's the open review list
|
||
(scan.md, category B). `scripts/check_corrections.py` verifies every
|
||
correction op against the actual files (a remove that matches nothing is
|
||
a silent no-op — it caught a mis-filed entry in CVSIX).
|
||
|
||
## Development
|
||
|
||
- **Run**: `./run.sh` → http://localhost:8599 (fixed port; no-ops if a
|
||
server is already running). The chart loads plotly.js from a CDN; for
|
||
fully offline use set `F_INLINE_PLOTLY=1` in `run.sh`.
|
||
- **Test**: `./run_tests.sh`
|
||
1. `tests/test_app.py` — app-level tests via Streamlit AppTest (no
|
||
browser). Memory: one data bundle is ~2.3 GB, so this process keeps
|
||
at most ONE AppTest alive (see its header comment).
|
||
2. `tests/test_e2e_browser.py` — Playwright + headless Chromium driving
|
||
the real page with real keystrokes; needs the server running on 8599.
|
||
One-time setup: `.venv/bin/pip install playwright` and
|
||
`.venv/bin/python -m playwright install chromium`.
|
||
- **Gotchas**
|
||
- Streamlit caches imported modules per process: **restart the server**
|
||
after editing any `.py` (kill the old one first — `run.sh` refuses to
|
||
double-start).
|
||
- `st.cache_data` caches the portfolio + tax simulations: they recompute
|
||
only when symbols/scheme/cost/tax rates change, not on window or curve
|
||
toggles.
|
||
- `settings.json` (gitignored) persists UI inputs across reloads and
|
||
restarts; delete it to reset. Saved portfolios live in `portfolios.json`.
|
||
- Data cache: `.cache/*.parquet`; rebuild via the sidebar checkbox
|
||
(first build ~1 min for ~4k symbols).
|
||
|
||
## Known simplifications (roadmap)
|
||
|
||
- No loss carryover or carryforward across years; no wash-sale rules.
|
||
- Distributed capital gains taxed entirely at the long-term rate.
|
||
- Single (federal-like) tax bracket; no state taxes, no AMT.
|
||
- Equal treatment of benchmark for beta/alpha (CAPM, rf = 0 by default).
|
||
|
||
Ideas: vectorbt sweeps over rebalance schemes, NiceGUI/Textual frontend,
|
||
empyrical-reloaded metrics, monthly (not yearly) loss netting, tax-loss
|
||
harvesting simulation.
|
||
|
||
## Local notes
|
||
|
||
`notes/tax.md` (gitignored) holds the personal tax profile, verified rate tables,
|
||
inherited-IRA rules/strategy, and the vehicle plan for this account. Update it when
|
||
tax facts change.
|