Commit Graph

9 Commits

Author SHA1 Message Date
3c5bee96ce Clean multi-vintage 'ctime' files: 1,065 files, 2.2M stale rows
Some dump files carry a trailing 'ctime' column: an older pipeline
appended a FULL re-download of the symbol per download (up to ~28
vintages per date, 2019-2022). The consumer keeps the last row per date,
so it silently used stale (~2021) vintages where Yahoo later revised
values. scripts/clean_multivintage.py keeps, per date, the rows from the
newest ctime (event files keep distinct same-day amounts; history/split
keep the single newest row), drops the ctime column, sorts by date.
Applied to the data root and overrides/frozen (which carried the same
artifact): 1,065 files, 2,202,712 stale rows dropped.

Also: stripped 3 UTF-8 BOMs (teg/lo/krft), refreshed the event backup
with the cleaned files, and re-ran the double-listing scan: category C
(83 symbols, 18,636 repeated within-file rows) is fully explained by the
multi-vintage artifact and is now gone; A (6,589 pairs, corrected) and
B (12,366 pairs, open review list) are unchanged. Whole-set structural
QC after cleaning: 0 duplicate dates, 3 syms/15 rows of OHLC invariant
violations, 15 syms/2,897 rows of non-positive prices (mostly
long-delisted tickers), 8 unsorted files (reader sorts).
2026-08-31 22:25:02 -04:00
e20b2e30cb Double-listing: layout-agnostic invariants, whole-dump sweep, bulk corrections
Yahoo changed the shape of its event feed between downloads: a 2026-08
re-download of CVIX/JLPSX shows it no longer returns capitalGain events at
all (the dividend stream still carries the year-end rows), while stale
tickers now return no events at all. File-specific remove ops therefore
break silently on the next re-download, so the double-listing fix is now
expressed as layout-agnostic invariants applied at bundle assembly
(idempotent, hold for full and incremental builds):

  dedup: [[date, amount]]      keep at most one copy of a same-date/
      same-amount cross-file pair (the capitalGain copy when both present)
  drop_capg_copy: [date]       the capitalGain row on that date is the
      spurious copy of the dividend row

- data.py: apply_invariants() at bundle assembly + pure dedupe_event_rows()
  shared with verify_official and tests (8 new test cases, 22 passing)
- JLPSX/CVSIX corrections rewritten with the invariants (2019-08-08 now
  keeps the dividend amount per the verified same-date pattern)
- scripts/scan_double_listing.py: whole-dump sweep -> reports/double_listing/
  6,421 symbols scanned: 1,218 with same-amount pairs (6,589), 1,631 with
  differing-amount pairs (12,366, reported only - not distinguishable from
  legitimate same-day div+capg without per-fund official data), 83 with
  repeated within-file rows (ingest keep-last already collapses them)
- scripts/apply_dedup_corrections.py: bulk 'dedup' corrections for the
  1,218 same-amount symbols (1,212 new files; the 6 verified funds keep
  their explicit, official-verified corrections)
- scripts/check_corrections.py: integrity check for every correction op
  against the actual (frozen) files - caught a mis-filed CVSIX entry
2026-08-31 21:00:36 -04:00
e15cb51390 Corrections: resolve the 7 stale mismatches via the Yahoo double-listing pattern
Scanning the verified funds for same-date div/capg pairs showed the Yahoo
double-listing mechanism in every one of the 7 'mismatch' stale funds, and
the official fiscal-year totals pin down which amount is true:

- identical-amount pairs (JLPSX, GDEUX, GSOUX, FAEVX, CVSIX, FZAGX): the
  year-end distribution is in both files; keep the capitalGain copy,
  remove the dividend copy (GDEUX: FY2021-08 1.73 = 1.625+0.115 and
  FY2022-08 0.39 = 0.348+0.036 after; JLPSX/GSOUX/FAEVX/CVSIX same)
- differing-amount pairs (OTCRX, SHXIX, FZAGX, FGIZX, CVSIX): the
  dividend-file amount is the true one in every officially-verified case
  (OTCRX FY2023 0.85 = 0.852488, SHXIX FY2023 1.01 = 1.0078, FZAGX all
  four FYs exact); the capitalGain row is the spurious copy

After corrections GDEUX/SHXIX/CVSIX/JLPSX verify 'ok' against their
filings. The same pattern is extended to these funds' pre-2021 history
(marked mechanism-inferred in the correction notes). Remaining residuals:
FAEVX and FGIZX are also missing regular quarterly dividend rows in the
Yahoo dump (official FY totals exceed local even after dedup) - needs the
funds' per-date distribution archives; FIKAX's official extraction is
ambiguous (systematic ~0.12/yr offset = wrong class table in the
500-fund consolidated Fidelity N-CSRS).
2026-08-31 19:50:18 -04:00
1f5720d7db verify_official: 497/497K forms, per-share layout, candidate re-ranking
- EFTS queries now include 497/497K (many fund families publish their
  per-fund highlights there, not in the consolidated N-CSR) and re-rank
  hits by registrant name match (ticker/brand words), newest first,
  capped at 2 filings per CIK
- new parse_per_share_blocks for the JPMorgan-style 'Per share operating
  performance' table (per-class value blocks; dashes = zero)
- parse_highlights now tolerates row labels split across table cells
  (modernized N-CSRS format, e.g. Calamos 2026)
- region finders: word-flexible name patterns (US vs U.S., class letters),
  self-validating per-share regions (a candidate block must match the
  local series, so a name mention in notes doesn't attribute another
  fund's tables in a combined 58 MB report)
- main() keeps the best result across candidate filings (N-CSRS vs 497
  can round differently) and stops early on 'ok'
- local_series applies the corrections overlay so corrected funds verify
  against their filing

Results: JLPSX and CVSIX now 'ok' (all bounded fiscal years agree with
the official filings); CVSIX also gets a 2023-12-21 0.510 capital-gain
correction. bnd/pmaix still ok (no regression).
2026-08-31 18:10:39 -04:00
6536c9903e Tier-3 verification: cross-check distributions against SEC filings
scripts/verify_official.py locates each fund's latest N-CSR/N-CSRS/10-K/10-Q
via EDGAR full-text search (full fund-name phrase first, then ticker +
name words, then bare ticker), extracts the fund's Financial Highlights
tables (both Vanguard-style and Victory-style layouts, calendar and
non-calendar fiscal years, M/D/YY and month-name headers), matches the
share class by per-share distribution series + NAV magnitude, and
compares per period window against the local Yahoo CSVs (frozen copies
for stale tickers). Per-symbol JSONs + SUMMARY.md land in
reports/xcheck_official/; results are cached per fund.

Run on the 29 curated funds: 8 ok (exact to 3dp, e.g. VTSAX 2021-2026H1),
1 mismatch (CVSIX FY2009: local 1.104 vs official 0.81 - the Yahoo
2008-12-18 row of 0.292 looks spurious), 3 weak-match (uncovered doc
formats, e.g. Leuthold), 17 not-found (mostly ETF families whose
10-K layouts aren't covered yet).
2026-08-31 14:41:03 -04:00
02aa750a06 Freeze stale (Yahoo-empty) tickers: snapshot final series in overrides/frozen/
230 tickers whose Yahoo chart responses now come back without a timestamp
array (terminated/merged funds): goget overwrites the .json on every pass
while ohlc.Conv skips the write, leaving the old CSVs as the last known
series. Snapshot them into overrides/frozen/ (git-tracked, audited in
reports/stale-funds.md) and make data.py prefer the frozen copies and
ignore any future data-root rewrite/delete for those symbols, so the
final series survives future goget runs. The cache manifest now covers
the overrides dir too; incremental refresh skips data-root files of
frozen symbols.
2026-08-31 13:23:24 -04:00
2bbe4e58ec Style-tilt battery + commentary in the fund report pipeline
fundlab/styletilt.py: 22 style/asset sleeves regressed on excess-of-T-bill
returns (full history + 5y); BIC forward selection identifies the tilt
stack; residual-alpha verdict ('factor exposure, not skill' when t<1.75);
data-driven English commentary with sign-specific phrasing. Rendered as a
'Style tilts' block (factor table + prose) in both the app and the HTML
report. All 24 pre-built + 8 ad-hoc fund reports rebuilt.
2026-08-30 20:36:04 -04:00
3fbf332b31 Per-fund report: app Summary page + narrative engine + mix_series beta fix
- fundlab/narrative.py: data-driven English prose per fund (performance,
  drivers tiered by fit, explicit 'what we do NOT know', bottom line)
- fundlab/reportdata.py: static build -> reports/report_data.json
- app.py Fund Lab Summary: at-a-glance table + per-fund expanders
  (narrative, equity curve, period table with fund-ref gap, drivers,
  reference mix, tax, cluster peers)
- fundlab/report.py: narrative in the HTML report; forward-selected
  reference (weak-fit funds anchor to cash); SLEEVE_DESC exposure
  explanations
- BUG: mix_series() never applied the betas (reference curves were raw
  sleeve sums; JLPSX 'reference' +407% vs fund +123%) - fixed and all
  reference curves/tables regenerated
- reports/fund_report.html + report_data.json regenerated
2026-08-30 17:36:58 -04:00
328855a926 Fund report: 11 candidates + 13 shortlist funds, self-contained HTML
fundlab/report.py -> reports/fund_report.html (20 MB, plotly inlined,
opens offline). Per fund: max-history equity curve (fund vs fitted
reference vs IVV); performance table (full/5y/1y, the 5 market
episodes, calendar years) with the fund-minus-reference period-alpha
column; drivers (reference-model R²/alpha/t + 34-sleeve signature +
curated decomposition verdict and N-PORT cross-check notes); the
reference mix explained sleeve-by-sleeve (what each exposure actually
is, plus net-cash/net-levered read); tax character + taxable/IRA
placement; and a peer table of the 4 best funds in the same k=30
return-driver cluster with computed advantages/disadvantages.

Weak-fit (R²<0.5) funds anchor their tables to CASH rather than the
statistically-thin forward-selected mix (which can be an offsetting
VIX/duration spec combination whose path is meaningless); the loadings
are still shown with a 'weak fit' caveat.
2026-08-30 17:00:41 -04:00