dedup/README.md
Greg Pomerantz 434330be1d Initial release: parallel duplicate finder with checkpointing
- SHA-256 hashing with multiprocessing (configurable workers)
- Resumable via checkpoints (saves every N files, auto-resumes)
- Reports exact matches, hash-only matches, same-name/different-content, unique
- --move (quarantine to __duplicates__/) or --delete
- Python 3.8+, stdlib only
2026-09-30 17:06:56 -04:00

78 lines
2.5 KiB
Markdown

# dedup
Find and remove duplicate files between two directories using SHA-256 hashing with parallel workers.
## Requirements
- Python 3.8+
- No external dependencies (standard library only)
## Usage
```bash
# Compare two directories (read-only, generates a report)
python3 dedup.py /path/to/backup /path/to/original
# Move duplicates to a __duplicates__/ quarantine folder
python3 dedup.py --move /path/to/backup /path/to/original
# Permanently delete duplicates
python3 dedup.py --delete /path/to/backup /path/to/original
# Control parallelism and checkpointing
python3 dedup.py --workers 8 --checkpoint-every 500 /large/backup /original
# Restrict to specific extensions
python3 dedup.py --ext .jpg,.png,.cr2 /path/to/backup /original
# Top-level files only (no recursion)
python3 dedup.py --no-subdirs /path/to/backup /original
# Remove checkpoint data and start fresh
python3 dedup.py --clean /path/to/backup /original
```
## Options
| Flag | Description |
|------|-------------|
| `--move` | Move duplicates to `__duplicates__/` folder (safe, reversible) |
| `--delete` | Permanently delete duplicates (irreversible) |
| `--no-subdirs` | Only compare top-level files |
| `--ext .jpg,.png` | Restrict to specific extensions (default: all media) |
| `--report PATH` | Write report to PATH (default: `dirA/duplicates_report.txt`) |
| `--workers N` | Number of parallel hash workers (default: CPU count) |
| `--checkpoint-every N` | Files between checkpoints, 0 to disable (default: 1000) |
| `--clean` | Remove checkpoint data and exit |
| `--version` | Show version |
## How It Works
1. **Phase 1** — Hash all files in the reference directory (dir B)
2. **Phase 2** — Hash all files in the source directory (dir A)
3. **Phase 3** — Compare hashes and generate a report
Files are classified as:
- **Exact matches** — same filename and same content hash
- **Hash matches** — same content, different filename
- **Same name, different content** — same filename, different content (not duplicates)
- **Unique to dirA** — no matching file in dirB
## Checkpointing
For large directories, progress is saved every `--checkpoint-every` files
(default: 1000) to `dirA/.dedup_checkpoint/`. If the process is interrupted,
re-run the same command to resume from where it left off.
Checkpoints are automatically cleaned after a successful run. Use `--clean`
to manually discard them.
## Output
A text report is written with sections for each match category, showing
file paths and SHA-256 hashes.
## License
MIT