Some dump files carry a trailing 'ctime' column: an older pipeline appended a FULL re-download of the symbol per download (up to ~28 vintages per date, 2019-2022). The consumer keeps the last row per date, so it silently used stale (~2021) vintages where Yahoo later revised values. scripts/clean_multivintage.py keeps, per date, the rows from the newest ctime (event files keep distinct same-day amounts; history/split keep the single newest row), drops the ctime column, sorts by date. Applied to the data root and overrides/frozen (which carried the same artifact): 1,065 files, 2,202,712 stale rows dropped. Also: stripped 3 UTF-8 BOMs (teg/lo/krft), refreshed the event backup with the cleaned files, and re-ran the double-listing scan: category C (83 symbols, 18,636 repeated within-file rows) is fully explained by the multi-vintage artifact and is now gone; A (6,589 pairs, corrected) and B (12,366 pairs, open review list) are unchanged. Whole-set structural QC after cleaning: 0 duplicate dates, 3 syms/15 rows of OHLC invariant violations, 15 syms/2,897 rows of non-positive prices (mostly long-delisted tickers), 8 unsorted files (reader sorts).
34 lines
535 B
Plaintext
34 lines
535 B
Plaintext
Date,Dividends
|
|
1999-11-10,0.0
|
|
2000-11-13,0.0
|
|
2001-12-21,0.29
|
|
2002-10-31,0.044
|
|
2010-12-02,0.0
|
|
2011-12-12,0.0
|
|
2012-12-13,0.052
|
|
2013-12-13,0.0
|
|
2014-12-12,0.0
|
|
2015-12-11,0.0
|
|
2016-12-06,0.061
|
|
2017-07-20,0.0
|
|
2017-10-12,0.07
|
|
2017-12-04,0.278
|
|
2018-04-11,0.319
|
|
2018-07-19,0.284
|
|
2018-10-11,0.27
|
|
2018-12-06,0.267
|
|
2019-04-10,0.243
|
|
2019-07-18,0.223
|
|
2019-10-10,0.318
|
|
2019-12-05,0.26
|
|
2020-04-08,0.0
|
|
2020-04-30,0.081
|
|
2020-06-30,0.089
|
|
2020-07-31,0.094
|
|
2020-08-31,0.124
|
|
2020-09-30,0.131
|
|
2020-10-30,0.1
|
|
2020-11-30,0.106
|
|
2020-12-31,0.128
|
|
2021-01-29,0.093
|