Find All Duplicate & Similar Files in Any Folder

Select a folder and Sakarto automatically groups every duplicate and similar file — images, videos, and audio — using your chosen algorithm. No reference file needed. Results appear live as the scan runs.

Images Videos Audio 100% Local Whole Folder at Once Live Results

How It Works

  1. 1. Pick a Folder
    Select any folder on your computer — subfolders are included automatically
  2. 2. Scan Runs
    Sakarto processes every file and builds fingerprints inside your browser. No data leaves your machine.
  3. 3. Groups Appear Live
    Duplicate groups appear as they are found — no waiting for the full scan to finish
  4. 4. Move or Delete
    Select files from any group and move, delete, or compare them side-by-side
Looking for copies of one specific file? Try Reverse Search instead — upload a reference file and Sakarto finds everything similar to it in a folder.

Image & Video Duplicate Finder — 7 Algorithms

Scan a folder and Sakarto groups all visually similar images and videos. Every algorithm produces a 64-bit hash compared using Hamming distance — lower threshold = stricter matching. Pick based on what kind of duplicates you expect.

Comparison of 7 image and video duplicate-detection algorithms
Algorithm How it works Canvas / Basis Rotation-safe Re-compression Brightness shifts Speed
Color Signature Compares YUV colour values across a 24×24 regional grid. The only algorithm that actually compares colour, not just luminance. 24×24 YUV grid ✓ Good ⚠ Partial ⚡ Fast
aHash Resizes to 16×16, converts to grayscale, compares each pixel to the overall mean. The simplest and fastest hash. 16×16 grayscale ✓ Good ⚠ Moderate ⚡ Fastest
BlockHash Divides a 64×64 canvas into 16×16 blocks, averages each block, compares to median. Block averaging absorbs JPEG noise and compression artifacts. 64×64, 16 blocks ✓ Excellent ✓ Good ⚡ Very Fast
dHash Resizes to 17×16 (one extra column), compares each pixel to the one on its right. Encodes horizontal brightness gradients, not absolute values. 17×16 gradient ✓ Good ✓ Excellent ⚡ Very Fast
pHash Resizes to 32×32, applies a 2D Discrete Cosine Transform, extracts top-left 8×8 low-frequency coefficients, thresholds against mean. Frequency-domain encoding is very stable across edits. 32×32 DCT ✓ Excellent ✓ Good ⚠ Moderate
wHash Resizes to 16×16, applies a Haar Wavelet Transform row-wise then column-wise, extracts 8×8 LL sub-band, thresholds against mean. Similar quality to pHash but faster. 16×16 Haar ✓ Good ✓ Good ⚡ Fast
ORB Uses OpenCV.js WASM to detect FAST keypoints and compute BRIEF descriptors. Matches keypoints spatially using Hamming distance + Lowe's ratio test. Threshold is inverted: higher = stricter. Keypoint features ✓ Good ✓ Good Slower — O(n²)
ORB threshold direction is inverted. For all other algorithms, lower threshold = stricter. For ORB, higher threshold = stricter (more matching keypoints required). ORB supports videos — 3 frames are extracted and the frame with the most keypoints is used. ORB requires OpenCV.js to load before scanning begins.
Color Signature

Color Signature Duplicate Finder

The only algorithm that compares actual colour — YUV values across a 24×24 regional grid. Finds copies sharing the same colour palette even when re-encoded or slightly edited. Colour-shifted versions of the same image will not match.

Best for: Photos with the same scene composition, re-compressed social media images, colour-similar duplicates where colour accuracy matters.
Fast — Web Worker + OffscreenCanvas
Find Duplicates
aHash

Average Hash (aHash) Duplicate Finder

The fastest duplicate finder. Resizes to 16×16, converts to grayscale, and compares each pixel to the overall mean — groups files sharing the same broad brightness distribution. Best for exact duplicates and resized copies.

Best for: Large libraries where speed matters most. Catches exact duplicates and resized copies instantly.
Fastest of all hash algorithms
Find Duplicates
BlockHash

Block Hash Duplicate Finder

Divides a 64×64 canvas into 16×16 blocks and averages each. Block averaging absorbs JPEG compression noise and image artifacts, making it more tolerant of degraded copies than pixel-level hash methods.

Best for: Noisy, heavily compressed, or low-quality copies. Scanned documents saved at different qualities.
Very Fast — Web Worker + OffscreenCanvas
Find Duplicates
dHash

Difference Hash (dHash) Duplicate Finder

Resizes to 17×16 (one extra column for gradients), converts to grayscale, and compares each pixel to the one on its right. Encodes relative brightness direction — not absolute values — making it robust to brightness and exposure changes.

Best for: Colour-graded or exposure-adjusted versions of the same photo. HDR vs standard edits. Brightness-shifted copies.
Very Fast — Web Worker + OffscreenCanvas
Find Duplicates
pHash

Perceptual Hash (pHash) Duplicate Finder

Resizes to 32×32, applies a separable 2D Discrete Cosine Transform, and extracts the top-left 8×8 low-frequency coefficients. Frequency-domain encoding is very stable across re-compression, mild colour grading, sharpening, and format conversion.

Best for: Watermarked copies, format conversions (JPEG→PNG→WebP), lightly edited variations. The best all-round hash for photo deduplication.
Fast — Web Worker + OffscreenCanvas
Find Duplicates
wHash

Wavelet Hash (wHash) Duplicate Finder

Resizes to 16×16, applies a Haar Wavelet Transform (averages and differences row-wise then column-wise), and extracts the 8×8 LL sub-band. Produces quality similar to pHash at lower CPU cost — the best speed/quality balance.

Best for: Large collections where speed matters but you want better quality than aHash. Images with local edits or regional modifications.
Fast — Web Worker + OffscreenCanvas
Find Duplicates
ORB

Feature Matching (ORB) Duplicate Finder

Uses OpenCV.js (WebAssembly) to detect FAST keypoints and compute BRIEF binary descriptors. Matches descriptors with Hamming distance + Lowe's ratio test. The only algorithm that handles rotation, perspective distortion, and cropping. Supports images and videos — 3 frames are extracted per video and the frame with the most keypoints is used.

Best for: Rotated, cropped, or perspective-transformed duplicates. When no other algorithm works. Note: threshold direction is reversed — higher = stricter.
Slower — O(n²) comparisons
Find Duplicates

📊 Which Image / Video Algorithm?

Quick reference for choosing based on what kind of duplicates you expect to find.

Which algorithm to use for different types of duplicate images and videos
I want to find… Color Sig. aHash BlockHash dHash pHash wHash ORB
Exact byte-for-byte copies Best
Resized versions (different resolution)
Re-compressed / different format Best
Brightness / exposure adjusted Best
Colour-graded (hue/saturation changed)
Same colour palette / scene Best
Watermarked or lightly edited Best
Rotated or mirrored copies Best
Cropped or perspective-warped Best
Very large folder (speed priority) BestAvoid
Videos as well as images ⚠ Best-keypoint frame

Audio Duplicate Finder — 3 Algorithms

Scan a folder and Sakarto groups all duplicate and similar audio files. Three algorithms, each measuring a different quality of sound. All audio processing continues in the background even when you switch browser tabs.

Which audio algorithm should I use? Chromaprint is the fastest and needs no external library — great for finding re-encoded copies of songs or podcasts. Essentia HPCP finds harmonically similar music like covers and live recordings. Meyda MFCC is best for voice, speech, and sound effects where timbral texture matters more than pitch.
Comparison of 3 audio duplicate-detection algorithms
Algorithm What it measures Best for Finds covers / alt versions Finds speech / podcasts Speed External library
Chromaprint
Spectral hash
Spectral band energy hashed into compact 15-bit frames. Captures overall frequency shape of the audio. Pure JavaScript, no CDN dependency. Lower threshold = stricter. Exact copies, re-encoded files, different bitrates of the same master track ⚠ Limited ✓ Good ⚡ Fastest — pure JS ✓ None
Essentia HPCP
Harmonic chroma
12-bin Harmonic Pitch Class Profile per frame via Essentia.js WASM at 44,100 Hz. Captures chord and pitch class content. Compares using cosine similarity. Higher threshold = stricter. Music with harmonic similarity — covers, transpositions, live vs studio recordings ✓ Excellent ✗ Poor ⚠ Moderate — WASM init 1–3s once ⚠ Essentia.js (CDN)
Meyda MFCC
Timbral texture
13 Mel-Frequency Cepstral Coefficients per frame via Meyda.js at 22,050 Hz. Analyses first 10 seconds only. Captures the perceptual "colour" of sound. Lower threshold = stricter. Podcasts, voice memos, sound effects — audio where timbre matters more than pitch content ✗ Poor ✓ Excellent ⚡ Fast — 10s window per file ⚠ Meyda.js (CDN)

⭐ What Users Say

Based on 187 reviews

"Found 3,000 duplicate photos in under 5 minutes. No install, no upload — my files never left my laptop."

"The Chromaprint audio finder is incredible. Cleaned up my entire podcast archive in one session."

"The ORB algorithm found rotated scans I had no idea were duplicates. Nothing else catches those."

"pHash found hundreds of re-compressed duplicates I'd been accumulating for years. Privacy is a huge plus too."


Frequently Asked Questions

How does the duplicate file finder work?
You select a folder on your computer. Sakarto scans every file inside it, builds a fingerprint for each one using your chosen algorithm, then compares all fingerprints to find similar or identical files. Everything runs inside your browser — no files are uploaded to any server.
Which algorithm should I use to find duplicate images?
Start with Color Signature for most photo collections. Use pHash or wHash for mixed or edited collections. Use aHash or BlockHash for speed on very large folders. Use ORB only for rotated or perspective-warped copies.
Which algorithm should I use to find duplicate audio files?
Use Chromaprint for re-encoded songs at different bitrates — it is the fastest and needs no CDN. Use Essentia HPCP for cover versions and live recordings. Use Meyda MFCC for podcasts, voice recordings, and sound effects.
What is the difference between Find Duplicates and Reverse Search?
Find Duplicates scans an entire folder and groups all similar files together automatically. Reverse Search lets you upload one specific reference file and finds everything in a folder that looks or sounds similar to that one file.
Does the duplicate finder work on Firefox or Safari?
Scanning and viewing results works in all modern browsers. However, the File System Access API needed for folder selection, file moving, and deletion is only available in Chrome and Edge 86+. Firefox and Safari fall back to a file picker that lets you select individual files instead of whole folders.
Can I delete files directly from the duplicate finder?
Yes. Select files from any duplicate group and click Delete. With Queue Mode on (default), files are staged for review before anything is permanently removed. Warning: deletion via the File System Access API bypasses the OS recycle bin. Deleted files cannot be recovered.