Some dump files carry a trailing 'ctime' column: an older pipeline appended a FULL re-download of the symbol per download (up to ~28 vintages per date, 2019-2022). The consumer keeps the last row per date, so it silently used stale (~2021) vintages where Yahoo later revised values. scripts/clean_multivintage.py keeps, per date, the rows from the newest ctime (event files keep distinct same-day amounts; history/split keep the single newest row), drops the ctime column, sorts by date. Applied to the data root and overrides/frozen (which carried the same artifact): 1,065 files, 2,202,712 stale rows dropped. Also: stripped 3 UTF-8 BOMs (teg/lo/krft), refreshed the event backup with the cleaned files, and re-ran the double-listing scan: category C (83 symbols, 18,636 repeated within-file rows) is fully explained by the multi-vintage artifact and is now gone; A (6,589 pairs, corrected) and B (12,366 pairs, open review list) are unchanged. Whole-set structural QC after cleaning: 0 duplicate dates, 3 syms/15 rows of OHLC invariant violations, 15 syms/2,897 rows of non-positive prices (mostly long-delisted tickers), 8 unsorted files (reader sorts).
486 B
486 B
| 1 | Date | Dividends |
|---|---|---|
| 2 | 2003-12-26 | 0.036 |
| 3 | 2005-12-23 | 0.306 |
| 4 | 2006-07-05 | 0.057 |
| 5 | 2006-12-21 | 0.959 |
| 6 | 2007-07-03 | 0.402 |
| 7 | 2007-12-20 | 2.66 |
| 8 | 2008-07-02 | 0.319 |
| 9 | 2008-12-19 | 0.719 |
| 10 | 2009-07-02 | 0.15 |
| 11 | 2009-12-18 | 0.031 |
| 12 | 2010-04-30 | 0.044 |
| 13 | 2010-07-07 | 0.028 |
| 14 | 2010-12-17 | 0.432 |
| 15 | 2011-07-07 | 0.073 |
| 16 | 2011-12-16 | 0.403 |
| 17 | 2012-12-17 | 0.62 |
| 18 | 2013-07-02 | 0.079 |
| 19 | 2013-12-16 | 0.813 |
| 20 | 2014-07-02 | 0.152 |
| 21 | 2014-12-16 | 1.32 |
| 22 | 2015-07-02 | 0.178 |
| 23 | 2015-12-16 | 1.092 |
| 24 | 2016-07-05 | 0.092 |
| 25 | 2016-12-16 | 0.818 |
| 26 | 2017-07-07 | 0.031 |
| 27 | 2017-12-18 | 1.125 |
| 28 | 2018-07-06 | 0.198 |
| 29 | 2018-12-17 | 1.15 |