Some dump files carry a trailing 'ctime' column: an older pipeline appended a FULL re-download of the symbol per download (up to ~28 vintages per date, 2019-2022). The consumer keeps the last row per date, so it silently used stale (~2021) vintages where Yahoo later revised values. scripts/clean_multivintage.py keeps, per date, the rows from the newest ctime (event files keep distinct same-day amounts; history/split keep the single newest row), drops the ctime column, sorts by date. Applied to the data root and overrides/frozen (which carried the same artifact): 1,065 files, 2,202,712 stale rows dropped. Also: stripped 3 UTF-8 BOMs (teg/lo/krft), refreshed the event backup with the cleaned files, and re-ran the double-listing scan: category C (83 symbols, 18,636 repeated within-file rows) is fully explained by the multi-vintage artifact and is now gone; A (6,589 pairs, corrected) and B (12,366 pairs, open review list) are unchanged. Whole-set structural QC after cleaning: 0 duplicate dates, 3 syms/15 rows of OHLC invariant violations, 15 syms/2,897 rows of non-positive prices (mostly long-delisted tickers), 8 unsorted files (reader sorts).
18 lines
280 B
Plaintext
18 lines
280 B
Plaintext
Date,Dividends
|
|
2015-08-12,0.263
|
|
2015-11-12,0.4125
|
|
2015-11-13,0.44
|
|
2016-02-12,0.46
|
|
2016-05-12,0.51
|
|
2016-08-11,0.525
|
|
2016-11-09,0.53
|
|
2017-02-13,0.535
|
|
2017-05-16,0.555
|
|
2017-08-11,0.57
|
|
2017-11-14,0.615
|
|
2018-02-14,0.62
|
|
2018-05-14,0.625
|
|
2018-08-14,0.63
|
|
2018-11-14,0.635
|
|
2019-02-14,0.64
|