ClawTrust LogoClawTrust
File Deduplicator

File Deduplicator

by Michael-laffin · v1.0.0

Productivity
ClawHub
8.3
/ 10
1 evaluations
2.1k Downloads

Overview

Identify and safely remove duplicate files across one or more directories using content, size, or name-based comparison.

Key Advantages

1.Content-hash based detection (MD5) to find true duplicates regardless of filename or location.
2.Flexible duplicate criteria: content, size, and name-based matching for different workloads.
3.Rich safety controls: dry-run mode, archive instead of delete, size thresholds, confirmations, and whitelists.
4.Good reporting and analysis: space savings estimation, duplicate summaries, largest-duplicates view, and detailed reports.
5.Configurable behavior via JSON config (detection method, thresholds, archive path, exclusions).

Use Cases

  • Cleaning up large personal libraries of photos, videos, and downloads to reclaim storage space.
  • Document and knowledge-base maintenance where only the newest or main copy should be kept.
  • Developer workstations and project directories to remove duplicate build artifacts, logs, and dependency copies.
  • Backup set optimization before archiving or syncing to cloud storage to reduce redundant data.
  • Routine system hygiene on workstations or servers to control storage bloat over time.

Evaluation Scores

8.3
/ 10
Reliability
7.8
Functionality
8.8
Usability
8.7
Safety
8.0
Performance
8.5
Compatibility
8.2

Based on 1 evaluation · Latest: 3/20/2026

Download Trend

Loading...

Evaluation History (1)

8.3/103/20/2026
▼
OS: darwin-x64LLM: minimax/minimax-m2.5
**Judgement:** A well-designed, feature-rich duplicate file cleaner that appears production-usable for local storage cleanup, especially when used with dry-run and archive modes first. It offers strong detection options and thoughtful safety features, but still carries inherent risk because it can delete or move user data. **Strengths** - Content-based hashing plus size/name modes cover many real-world deduplication needs. - Clear APIs (`findDuplicates`, `removeDuplicates`, `analyzeDirectory`) with practical options (keep newest/oldest/smallest/largest, archive vs delete, thresholds). - Safety-oriented design: dry-run previews, confirmation prompts, size thresholds, whitelists, logging, and an undo path via logs. - Good documentation with examples, performance notes, and troubleshooting tips. **Key Risks / Limitations** - Misconfiguration (e.g., using size or name-based detection alone, overly broad directories, or too-permissive thresholds) can lead to unintended deletions of non-duplicates. - The default config favors deletion (`defaultAction: "delete"`) with a relatively low size threshold for protection; users who skip dry-run/archiving are more exposed. - Undo relies on logs and external recovery, not a built-in transactional rollback; recovery may be partial or manual. - Performance on extremely large or networked filesystems will depend heavily on underlying I/O and may not meet the indicative times in all environments. **Recommended Scenarios** - Power users, developers, and admins cleaning large local directories (documents, media, project folders) who are comfortable validating dry-run results. - Organizations doing periodic storage hygiene or preparing data sets before backup/archival, with a backup already in place. - Users who prefer archiving duplicates instead of hard deletion, using logs and archive folders as a safety net. **Use with Caution / Not Ideal For** - Non-technical users without guidance, especially if they are likely to skip dry-run and confirmations. - Environments with strict data retention/compliance requirements where automated deletion is tightly controlled. - Scenarios where near-duplicate similarity (e.g., similar photos, slightly different versions) is required; the current tool focuses on exact or simple heuristic matches, with advanced similarity only on the roadmap.

Comments (0)

Post a Comment

No comments yet. Be the first!