Some dump files carry a trailing 'ctime' column: an older pipeline appended a FULL re-download of the symbol per download (up to ~28 vintages per date, 2019-2022). The consumer keeps the last row per date, so it silently used stale (~2021) vintages where Yahoo later revised values. scripts/clean_multivintage.py keeps, per date, the rows from the newest ctime (event files keep distinct same-day amounts; history/split keep the single newest row), drops the ctime column, sorts by date. Applied to the data root and overrides/frozen (which carried the same artifact): 1,065 files, 2,202,712 stale rows dropped. Also: stripped 3 UTF-8 BOMs (teg/lo/krft), refreshed the event backup with the cleaned files, and re-ran the double-listing scan: category C (83 symbols, 18,636 repeated within-file rows) is fully explained by the multi-vintage artifact and is now gone; A (6,589 pairs, corrected) and B (12,366 pairs, open review list) are unchanged. Whole-set structural QC after cleaning: 0 duplicate dates, 3 syms/15 rows of OHLC invariant violations, 15 syms/2,897 rows of non-positive prices (mostly long-delisted tickers), 8 unsorted files (reader sorts).
25 lines
400 B
Plaintext
25 lines
400 B
Plaintext
Date,Dividends
|
|
2009-06-22,0.618
|
|
2009-12-21,0.04
|
|
2010-03-29,0.041
|
|
2010-06-28,0.404
|
|
2010-09-20,0.021
|
|
2010-12-22,0.022
|
|
2011-03-21,0.058
|
|
2011-06-22,0.765
|
|
2012-03-26,0.16
|
|
2012-06-25,0.5
|
|
2013-03-22,0.214
|
|
2013-06-24,0.348
|
|
2014-03-24,0.369
|
|
2014-06-23,0.167
|
|
2014-09-22,0.098
|
|
2014-12-19,0.357
|
|
2015-03-23,0.115
|
|
2015-06-22,0.625
|
|
2015-12-21,0.034
|
|
2016-06-20,0.665
|
|
2016-12-23,0.005
|
|
2017-03-27,0.1
|
|
2017-06-26,0.235
|