python
31 lines · 7 steps
Finding duplicate files by size then hash
A two-pass scan that groups files by size, then confirms duplicates by hashing only the candidates.
Explained by
highlit
1import hashlib
2from collections import defaultdict
3from pathlib import Path
4
5
6def _hash_file(path: Path, chunk_size: int = 65536) -> str:
7 digest = hashlib.sha256()
8 with path.open("rb") as fh:
9 for chunk in iter(lambda: fh.read(chunk_size), b""):
10 digest.update(chunk)
11 return digest.hexdigest()
12
13
14def find_duplicate_files(root: str | Path) -> dict[str, list[Path]]:
15 root = Path(root)
16 by_size: dict[int, list[Path]] = defaultdict(list)
17 for path in root.rglob("*"):
18 if path.is_file():
19 by_size[path.stat().st_size].append(path)
20
21 duplicates: dict[str, list[Path]] = defaultdict(list)
22 for size, paths in by_size.items():
23 if len(paths) < 2:
24 continue
25 for path in paths:
26 try:
27 duplicates[_hash_file(path)].append(path)
28 except OSError:
29 continue
30
31 return {digest: paths for digest, paths in duplicates.items() if len(paths) > 1}
01 / 01
STEP 01
‹ swipe to step through ›
Walkthrough
Space play
←→ step
click any line
Three takeaways
- 1Cheap filters first: grouping by size avoids hashing files that can't possibly match.
- 2Reading in fixed chunks keeps memory flat regardless of file size.
- 3Guarding I/O with try/except lets a scan survive unreadable files instead of crashing.
Related explainers
python
from fastapi import FastAPI, WebSocket, WebSocketDisconnect app = FastAPI()
Building a WebSocket chat with FastAPI
websockets
broadcast
connection-management
Intermediate
9 steps
python
import time import uuid from django.utils.deprecation import MiddlewareMixin
Attaching per-request context in Django
middleware
request lifecycle
multi-tenancy
Intermediate
7 steps
ruby
class LogAggregator BUCKET_FORMAT = "%Y-%m-%dT%H:%M" def initialize(entries)
Bucketing log entries by the minute in Ruby
aggregation
hashing
enumerable
Intermediate
5 steps
python
import random from typing import Iterator, List
How reservoir sampling picks k items
reservoir-sampling
streaming
randomness
Intermediate
5 steps
go
package logging import ( "context"
Deduplicating log attributes in Go's slog
decorator-pattern
structured-logging
immutability
Intermediate
8 steps
ruby
class ApplicationController < ActionController::Base EXPERIMENTS = { checkout_button_color: %w[control blue green], onboarding_flow: %w[control streamlined]
How A/B test cohorts are assigned in Rails
a-b-testing
cookies
hashing
Intermediate
8 steps
Share this explainer
Here's the card — post it anywhere.
Made with highlit — turn any snippet into a walkthrough like this in about a minute.
Explain your code
Embed this explainer
Drop the interactive walkthrough into a blog or docs. Views never cost a credit.
<iframe src="https://highlit.co/explainers/finding-duplicate-files-by-size-then-hash-explained-python-8c1d/embed?autoplay=1" width="100%" height="520" loading="lazy" style="border:0"></iframe>
Autoplay is on by default — add ?autoplay=0 to start paused.