dedup/README.md
Greg Pomerantz 434330be1d Initial release: parallel duplicate finder with checkpointing
- SHA-256 hashing with multiprocessing (configurable workers)
- Resumable via checkpoints (saves every N files, auto-resumes)
- Reports exact matches, hash-only matches, same-name/different-content, unique
- --move (quarantine to __duplicates__/) or --delete
- Python 3.8+, stdlib only
2026-09-30 17:06:56 -04:00

2.5 KiB

dedup

Find and remove duplicate files between two directories using SHA-256 hashing with parallel workers.

Requirements

  • Python 3.8+
  • No external dependencies (standard library only)

Usage

# Compare two directories (read-only, generates a report)
python3 dedup.py /path/to/backup /path/to/original

# Move duplicates to a __duplicates__/ quarantine folder
python3 dedup.py --move /path/to/backup /path/to/original

# Permanently delete duplicates
python3 dedup.py --delete /path/to/backup /path/to/original

# Control parallelism and checkpointing
python3 dedup.py --workers 8 --checkpoint-every 500 /large/backup /original

# Restrict to specific extensions
python3 dedup.py --ext .jpg,.png,.cr2 /path/to/backup /original

# Top-level files only (no recursion)
python3 dedup.py --no-subdirs /path/to/backup /original

# Remove checkpoint data and start fresh
python3 dedup.py --clean /path/to/backup /original

Options

Flag Description
--move Move duplicates to __duplicates__/ folder (safe, reversible)
--delete Permanently delete duplicates (irreversible)
--no-subdirs Only compare top-level files
--ext .jpg,.png Restrict to specific extensions (default: all media)
--report PATH Write report to PATH (default: dirA/duplicates_report.txt)
--workers N Number of parallel hash workers (default: CPU count)
--checkpoint-every N Files between checkpoints, 0 to disable (default: 1000)
--clean Remove checkpoint data and exit
--version Show version

How It Works

  1. Phase 1 — Hash all files in the reference directory (dir B)
  2. Phase 2 — Hash all files in the source directory (dir A)
  3. Phase 3 — Compare hashes and generate a report

Files are classified as:

  • Exact matches — same filename and same content hash
  • Hash matches — same content, different filename
  • Same name, different content — same filename, different content (not duplicates)
  • Unique to dirA — no matching file in dirB

Checkpointing

For large directories, progress is saved every --checkpoint-every files (default: 1000) to dirA/.dedup_checkpoint/. If the process is interrupted, re-run the same command to resume from where it left off.

Checkpoints are automatically cleaned after a successful run. Use --clean to manually discard them.

Output

A text report is written with sections for each match category, showing file paths and SHA-256 hashes.

License

MIT