Review directories recursively for duplicate files.
- SHA-256 hashing with multiprocessing (configurable workers) - Resumable via checkpoints (saves every N files, auto-resumes) - Reports exact matches, hash-only matches, same-name/different-content, unique - --move (quarantine to __duplicates__/) or --delete - Python 3.8+, stdlib only |
||
|---|---|---|
| .gitignore | ||
| dedup.py | ||
| LICENSE | ||
| README.md | ||
dedup
Find and remove duplicate files between two directories using SHA-256 hashing with parallel workers.
Requirements
- Python 3.8+
- No external dependencies (standard library only)
Usage
# Compare two directories (read-only, generates a report)
python3 dedup.py /path/to/backup /path/to/original
# Move duplicates to a __duplicates__/ quarantine folder
python3 dedup.py --move /path/to/backup /path/to/original
# Permanently delete duplicates
python3 dedup.py --delete /path/to/backup /path/to/original
# Control parallelism and checkpointing
python3 dedup.py --workers 8 --checkpoint-every 500 /large/backup /original
# Restrict to specific extensions
python3 dedup.py --ext .jpg,.png,.cr2 /path/to/backup /original
# Top-level files only (no recursion)
python3 dedup.py --no-subdirs /path/to/backup /original
# Remove checkpoint data and start fresh
python3 dedup.py --clean /path/to/backup /original
Options
| Flag | Description |
|---|---|
--move |
Move duplicates to __duplicates__/ folder (safe, reversible) |
--delete |
Permanently delete duplicates (irreversible) |
--no-subdirs |
Only compare top-level files |
--ext .jpg,.png |
Restrict to specific extensions (default: all media) |
--report PATH |
Write report to PATH (default: dirA/duplicates_report.txt) |
--workers N |
Number of parallel hash workers (default: CPU count) |
--checkpoint-every N |
Files between checkpoints, 0 to disable (default: 1000) |
--clean |
Remove checkpoint data and exit |
--version |
Show version |
How It Works
- Phase 1 — Hash all files in the reference directory (dir B)
- Phase 2 — Hash all files in the source directory (dir A)
- Phase 3 — Compare hashes and generate a report
Files are classified as:
- Exact matches — same filename and same content hash
- Hash matches — same content, different filename
- Same name, different content — same filename, different content (not duplicates)
- Unique to dirA — no matching file in dirB
Checkpointing
For large directories, progress is saved every --checkpoint-every files
(default: 1000) to dirA/.dedup_checkpoint/. If the process is interrupted,
re-run the same command to resume from where it left off.
Checkpoints are automatically cleaned after a successful run. Use --clean
to manually discard them.
Output
A text report is written with sections for each match category, showing file paths and SHA-256 hashes.
License
MIT