Financial analysis tools
Go to file
2026-09-01 21:42:49 -04:00
fundlab Style-tilt battery + commentary in the fund report pipeline 2026-08-30 20:36:04 -04:00
overrides Fix after-tax model: raw prices + reinvestment (was 2.7x over-taxing) 2026-08-31 23:58:15 -04:00
reports Fix Yahoo total-distribution double count (1,957 syms) 2026-09-01 16:23:03 -04:00
scripts Fix Yahoo total-distribution double count (1,957 syms) 2026-09-01 16:23:03 -04:00
tests Tax engine: apply NIIT + state/local to the equity path (not just recorded taxes) 2026-09-01 17:22:04 -04:00
.gitignore gitignore notes/ (personal tax profile); README pointer 2026-09-01 12:42:25 -04:00
adx-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
aef-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
app.py Add NIIT + state/local rates to the after-tax model (NYC support) 2026-09-01 09:37:58 -04:00
asa-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
brw-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
bto-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
chart_widget.py Stock & Portfolio Analyzer: full UI rework 2026-08-24 16:05:27 -04:00
clm-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
crf-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
data.py Fix Yahoo total-distribution double count (1,957 syms) 2026-09-01 16:23:03 -04:00
evg-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
families.py Stock & Portfolio Analyzer: full UI rework 2026-08-24 16:05:27 -04:00
fxby-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
grf-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
herz-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
iaf-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
kf-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
mci-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
metrics.py app: background cache refresh, per-benchmark stats, correlation tab, global date range 2026-08-25 18:15:47 -04:00
mxf-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
nro-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
peo-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
portfolio.py Stock & Portfolio Analyzer: full UI rework 2026-08-24 16:05:27 -04:00
portfolios.json Add LCRIX proxy portfolio (Leuthold Core, from 2026-03-31 N-CSRS asset mix) 2026-09-01 21:42:49 -04:00
portfolios.py Stock & Portfolio Analyzer: full UI rework 2026-08-24 16:05:27 -04:00
README.md gitignore notes/ (personal tax profile); README pointer 2026-09-01 12:42:25 -04:00
requirements.txt Stock & Portfolio Analyzer: full UI rework 2026-08-24 16:05:27 -04:00
run_tests.sh Fix after-tax model: raw prices + reinvestment (was 2.7x over-taxing) 2026-08-31 23:58:15 -04:00
run.sh Stock & Portfolio Analyzer: full UI rework 2026-08-24 16:05:27 -04:00
rvt-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
saba-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
swz-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
tax.py Tax engine: apply NIIT + state/local to the equity path (not just recorded taxes) 2026-09-01 17:22:04 -04:00
tyg-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
utf-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
vlt-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00
ztr-split.csv CEF pass: universe (SEC report, 973 -> 295 listed) + stage 1 screen + stage 2a character 2026-08-27 22:24:12 -04:00

Stock & Portfolio Analyzer

Interactive tool for analyzing individual securities and portfolios against local Yahoo Finance dumps (~/prog/fin/stocks, ~4k symbols).

Quick start

python3 -m venv .venv
.venv/bin/pip install -r requirements.txt
./run.sh    # serves the UI on the fixed port 8599 (http://localhost:8599)

First run builds a parquet cache in .cache/ (~1 min for 4k symbols); later runs load in well under a second. The cache tracks the data dir per-file (mtime + size in .cache/manifest.json), so when the download is updated, only the changed/added/removed symbols are re-read — a partial refresh takes seconds instead of a full ~1 min rebuild.

Modules

Module Purpose
data.py Ingest {sym}-history/dividend/capitalGain.csv -> cached parquet panels (date x symbol), with manifest-based incremental refresh when the data dir changes. Adj Close already includes distributions, so it drives pre-tax total returns.
metrics.py Total/annualized return, vol, Sharpe, Sortino, max drawdown, Calmar, CAPM beta/alpha. Pure pandas, all transparent.
portfolio.py Weighted portfolios with drift and periodic rebalancing to target weights (1W/1ME/QE/YE), one-way cost in bps. Spec grammar: commas join the elements of ONE portfolio (SYM or SYM:w, bare = equal weight), spaces separate DISTINCT symbols/portfolios (parse_items).
tax.py Simplified DAS after-tax engine: FIFO lots, 365-day long/short split, separate LT/ST/dividend rates plus federal NIIT and a state+local marginal rate (applied at ordinary rates — state/local have no preferential cap-gain rate). Headline curve = what you keep if you sell everything today (unrealized gains taxed daily by lot age).
chart_widget.py Self-contained plotly.js chart in an iframe: mouse zoom/pan, x clamped to the data, view edges snapped to first/last data points with day-precise labels, y tight-fit, every line re-based to 1.0 at the left edge.
portfolios.py Saved portfolio definitions in portfolios.json (name, spec, scheme, cost).
settings.json Persisted UI inputs (symbol/benchmark specs, scheme, costs, tax rates, period, curve/window mode) — restored on every page load and server restart; delete to reset.
app.py Streamlit UI: single "symbol or portfolio" spec field (page updates as soon as the input is valid; unknown symbols get click-to-fix "did you mean" suggestions) + a benchmark box with the same grammar (one benchmark per line; a line is a single symbol or a comma-joined portfolio, simulated with the same scheme/cost/tax rules — pre- and after-tax curves, first one drives beta/alpha), scheme/costs/tax rates, save + load/compare/delete portfolios (overlaid pre/after-tax curves), curve toggle (both / pre-tax only / after-tax only), stats table, allocation, per-year tax detail.

Data verification (Tier 3: against official filings)

Notes on the data itself:

  • Capital gains in this dump are mutual funds only. The dividend/capitalGain event files exist only for open-end mutual funds (spot-checked: no ETFs, CEFs, BDCs or individual stocks carry them — CEF/BDC distribution components are lumped into their dividend stream by Yahoo). For total returns only the distribution TOTAL matters, which is what Yahoo provides and what the corrections above fix.
  • There is no long-term/short-term split anywhere in the Yahoo data (an event is just a per-share amount), and the corrections never used one either — the official checks compare total distributions. tax.py therefore taxes capital-gain distributions at the long-term rate (lt_rate), a simplification: fund cap-gain distributions are predominantly long-term, but the per-fund LT/ST split would come from the fund company's annual tax statement (1099-DIV detail: boxes 2a/2b), the shareholder-report body, or a commercial feed (Lipper/Morningstar).
  • NYC residents: add the state+local layer. The sidebar's two extra rate fields exist for this. Set NIIT % to 3.8 if your MAGI is over the threshold, and State + local % to your NY+NYC MARGINAL rate sum (2025 single, from the Form IT-201 rate schedules): NYC is 3.876% above $50k of city taxable income; NY is 6.85% at $215,400-$1.077M of state taxable income and 9.65% at $1.077M-$5M. So a typical NYC household pays, on a capital-gain distribution, roughly 20% federal + 3.8% NIIT + 6.85-9.65% NY + 3.876% NYC = ~34-37% — which is exactly why high-distribution open-end funds lose so much more to tax than ETFs for NYC residents (see the per-year tax detail tab). Note: NY/NYC tax capital gains at ORDINARY rates (no preferential cap-gain rate), which the single State + local % field models correctly; the federal Long-term gains % field stays the preferential 0/15/20%.

~/prog/fin/stocks is a Yahoo dump and gets re-downloaded (overwritten), so fixes must live outside it. Pipeline:

  1. scripts/audit_stale.py — finds tickers whose latest history row is old and whose fresh Yahoo download is empty (delisted/merged funds, tickers Yahoo no longer serves). Snapshots their final {history,dividend,capitalGain}.csv + longName into overrides/frozen/{SYM}.{ext|json}. data.py shadows the data root with the frozen copies, so a re-download can't clobber them. 1b. overrides/event-backup/ — backup of every dividend/capitalGain file (refresh with scripts/backup_events.py after each dump update). Yahoo changed its event feed in 2026: for some funds it no longer returns capitalGains events at all, and terminated funds return no events, so a re-download can empty or orphan good event history. The data-root file wins while it is populated; an emptied/missing one falls back to the backup (data.py event_file). The ohlc converter likewise refuses to overwrite a populated event CSV with an empty download. (Yahoo omits the whole capitalGains JSON key when there are no events, so plain re-downloads in place usually leave old files untouched — the backup is the second line of defense.)
  2. scripts/verify_official.py [syms | --stale] — finds the fund's shareholder report (EDGAR EFTS for "TICKER", forms N-CSR/N-CSRS/ N-14/N-2/497) and parses the per-class Financial Highlights (or JPMorgan-style Per share operating performance) tables: period-by- period distribution totals compared against local dividends + capital-gains over the same windows, plus a spot check of the NAV-per-share row against the local close. Verdicts per symbol in reports/xcheck_official/{sym}.json: ok (all bounded fiscal years agree), mismatch (class matched, some year off — the report tells you which), weak-match (best class fit too poor to trust), not-found (fund not in any candidate filing). The local side applies the corrections overlay, so a corrected fund verifies against its filing. When several filings parse (e.g. the 497 annual and the N-CSRS, which can round differently), the best match wins.
  3. Confirmed findings go into overrides/corrections/{SYM}.json as auditable deltas (remove/replace/add of distribution rows, each entry dated and valued — amounts are what Yahoo reports, not the official filing's). data.py applies them on cache build. Examples:
    • CVSIX — 2008-12-18 0.292 duplicate of the same day's 0.641 (official FY09 = 0.81 balances without it); 2023-12-21 0.510 capital-gain row duplicated next to the day's 0.691 dividend (official FY2024-10 = 0.79 balances without it).
    • JLPSX — 13 year-end capital-gain distributions duplicated into the dividend file (FY20212025 all match the JPMorgan 497 after removal).

Remaining known data gaps (need the fund company's per-date distribution archive; fiscal-year totals alone can't reconstruct them): FAEVX and FGIZX — the Yahoo dump is missing the funds' regular quarterly dividend rows (official fiscal-year totals exceed local even after the double-listing dedup above); FIKAX — the official extraction is ambiguous (the 500-fund consolidated Fidelity N-CSRS has multiple near-identical sub-fund tables, and the best match is a systematic ~0.12/yr offset, i.e. the wrong share class).

Whole-dump double-listing sweep. scripts/scan_double_listing.py scans every symbol's effective event files for the pattern and writes reports/double_listing/scan.{md,json}. On the 2026-08 dump (6,421 symbols with event files): 1,218 symbols had same-date/same-amount cross-file pairs (6,589 pairs, 80% in December / fiscal year-end), 1,631 had same-date differing-amount pairs (12,366), and 83 had repeated same-date rows within one file (Yahoo repeating a row up to ~35x; the ingest already keeps the last). The same-amount pairs were corrected in bulk by scripts/apply_dedup_corrections.py (1,212 correction files, dedup invariant, mechanism-inferred and individually revertible). The differing-amount pairs are reported but NOT auto-corrected: without a per-fund official schedule they can't be distinguished from a legitimate same-day dividend + capital-gain pairing — that's the open review list (scan.md, category B). scripts/check_corrections.py verifies every correction op against the actual files (a remove that matches nothing is a silent no-op — it caught a mis-filed entry in CVSIX).

Development

  • Run: ./run.shhttp://localhost:8599 (fixed port; no-ops if a server is already running). The chart loads plotly.js from a CDN; for fully offline use set F_INLINE_PLOTLY=1 in run.sh.
  • Test: ./run_tests.sh
    1. tests/test_app.py — app-level tests via Streamlit AppTest (no browser). Memory: one data bundle is ~2.3 GB, so this process keeps at most ONE AppTest alive (see its header comment).
    2. tests/test_e2e_browser.py — Playwright + headless Chromium driving the real page with real keystrokes; needs the server running on 8599. One-time setup: .venv/bin/pip install playwright and .venv/bin/python -m playwright install chromium.
  • Gotchas
    • Streamlit caches imported modules per process: restart the server after editing any .py (kill the old one first — run.sh refuses to double-start).
    • st.cache_data caches the portfolio + tax simulations: they recompute only when symbols/scheme/cost/tax rates change, not on window or curve toggles.
    • settings.json (gitignored) persists UI inputs across reloads and restarts; delete it to reset. Saved portfolios live in portfolios.json.
    • Data cache: .cache/*.parquet; rebuild via the sidebar checkbox (first build ~1 min for ~4k symbols).

Known simplifications (roadmap)

  • No loss carryover or carryforward across years; no wash-sale rules.
  • Distributed capital gains taxed entirely at the long-term rate.
  • Single (federal-like) tax bracket; no state taxes, no AMT.
  • Equal treatment of benchmark for beta/alpha (CAPM, rf = 0 by default).

Ideas: vectorbt sweeps over rebalance schemes, NiceGUI/Textual frontend, empyrical-reloaded metrics, monthly (not yearly) loss netting, tax-loss harvesting simulation.

Local notes

notes/tax.md (gitignored) holds the personal tax profile, verified rate tables, inherited-IRA rules/strategy, and the vehicle plan for this account. Update it when tax facts change.