f/overrides/event-backup/hunax-dividend.csv
Greg Pomerantz 3c5bee96ce Clean multi-vintage 'ctime' files: 1,065 files, 2.2M stale rows
Some dump files carry a trailing 'ctime' column: an older pipeline
appended a FULL re-download of the symbol per download (up to ~28
vintages per date, 2019-2022). The consumer keeps the last row per date,
so it silently used stale (~2021) vintages where Yahoo later revised
values. scripts/clean_multivintage.py keeps, per date, the rows from the
newest ctime (event files keep distinct same-day amounts; history/split
keep the single newest row), drops the ctime column, sorts by date.
Applied to the data root and overrides/frozen (which carried the same
artifact): 1,065 files, 2,202,712 stale rows dropped.

Also: stripped 3 UTF-8 BOMs (teg/lo/krft), refreshed the event backup
with the cleaned files, and re-ran the double-listing scan: category C
(83 symbols, 18,636 repeated within-file rows) is fully explained by the
multi-vintage artifact and is now gone; A (6,589 pairs, corrected) and
B (12,366 pairs, open review list) are unchanged. Whole-set structural
QC after cleaning: 0 duplicate dates, 3 syms/15 rows of OHLC invariant
violations, 15 syms/2,897 rows of non-positive prices (mostly
long-delisted tickers), 8 unsorted files (reader sorts).
2026-08-31 22:25:02 -04:00

639 B

1DateDividends
22014-01-300.01
32014-02-270.01
42014-03-280.016
52014-04-290.025
62014-05-290.021
72014-06-270.016
82014-07-300.016
92014-08-280.012
102014-09-290.008
112014-10-300.012
122014-11-260.011
132014-12-050.4
142014-12-300.012
152015-01-290.011
162015-02-260.009
172015-03-300.004
182015-04-290.007
192015-05-280.011
202015-06-290.012
212015-07-300.015
222015-08-280.01
232015-09-290.004
242015-10-290.007
252015-11-270.006
262015-12-040.359
272015-12-300.005
282016-01-260.006
292016-02-240.006
302016-03-280.005
312016-04-260.007
322016-05-250.011
332016-06-270.012
342016-07-260.013
352016-08-260.011
362016-09-270.005
372016-10-260.006
382016-11-250.007