ClawTrust LogoClawTrust
Data Anomaly Detector

Data Anomaly Detector

by datadrivenconstruction · v1.0.0

Productivity
ClawHub
7.6
/ 10
1 evaluations
1.7k Downloads

Overview

Statistical anomaly detection for construction cost, schedule, productivity, and data quality in tabular datasets using pandas-based methods.

Key Advantages

1.Focused on construction domain (costs, schedules, productivity, invoice/sequences) rather than generic anomaly detection.
2.Multiple statistical techniques implemented (IQR, group Z-score, modified Z-score, rolling time-series Z-score, rule-based checks).
3.Covers several anomaly classes: outliers, trend deviations, impossible values, duplicates, and missing sequences.
4.Provides a structured anomaly model (dataclasses) and an aggregated AnomalyReport object for downstream processing.
5.Built-in markdown report generation for quick human review, including severity summaries and critical-anomaly focus.

Use Cases

  • Pre-flight QA/QC on project cost exports (e.g., Excel/CSV) to flag suspicious costs, negatives, and group-wise outliers before budgeting reviews.
  • Schedule health checks to detect impossible dates, excessively long activities, and non-milestone zero-duration tasks in CPM or Gantt exports.
  • Field productivity monitoring from timesheets and production quantities to identify abnormally high/low productivity for investigation.
  • Data pipeline validation for construction ERPs to catch duplicate records, missing invoice/PO sequences, and obvious data-entry errors.
  • Periodic anomaly scans on historical project data to support auditing, fraud detection, and lessons-learned analyses.

Evaluation Scores

7.6
/ 10
Reliability
6.6
Functionality
7.4
Usability
7.8
Safety
8.5
Performance
7.5
Compatibility
8.0

Based on 1 evaluation · Latest: 3/20/2026

Download Trend

Loading...

Evaluation History (1)

7.6/103/20/2026
▼
OS: linux-x64LLM: z-ai/glm-4.5-air
**Quick judgement** Solid, construction-focused anomaly detector for tabular data built on pandas/numpy/scipy. It offers a good breadth of statistical and rule-based checks (cost, schedule, productivity, duplicates, sequences), but has some edge-case and robustness gaps that users should be aware of when integrating into production pipelines. **Strengths** - Covers several key construction data domains: costs, schedules, productivity, duplicates, and numeric sequences. - Uses established statistical methods (IQR, Z-score, modified Z-score, rolling Z-score) plus business-rule checks (negative costs, impossible dates, zero-duration non-milestones, missing sequences). - Provides a clean data model (Anomaly, AnomalyReport) and a convenient `run_full_detection` orchestration method. - Markdown report generator produces human-readable summaries with severity breakdowns and a table of anomalies. **Key risks / limitations** - **Edge-case robustness:** Some methods may fail or behave poorly when data is degenerate: - Group-based Z-scores in `detect_cost_anomalies` can divide by zero when group std is 0, yielding infinities and potentially excessive false positives. - Productivity detection (`detect_productivity_anomalies`) can divide by zero when MAD is 0, leading to NaNs/inf and unpredictable behavior. - Sequence gap detection (`detect_sequence_gaps`) assumes at least one numeric sequence value; if all values are non-numeric or missing, `min()/max()` will error. - **Config/column assumptions:** `run_full_detection` does not fully guard against missing columns (e.g., a `sequence_column` or `key_columns` specified in config but absent in the DataFrame can raise exceptions). - **In-place DataFrame mutation:** Several methods add/overwrite columns (e.g., `duration`, `productivity`, `seq_num`, converted date types) directly on the provided DataFrame, which may surprise callers that expect inputs to remain unchanged. - **Domain thresholds not fully leveraged:** The defined construction-specific threshold dictionaries are not consistently used across detection methods, limiting domain-tuning out of the box. - **Time-series detector not wired into `run_full_detection`:** `detect_time_series_anomalies` must be called manually; it’s not integrated into the main pipeline. **Recommended usage scenarios** - Offline or batch QA/QC on construction project exports (cost/schedule/productivity tables) before reporting, forecasting, or executive reviews. - As a data quality gate in ETL pipelines for construction ERPs, cost management systems, or scheduling tools, with appropriate try/except wrapping and input validation. - Exploratory anomaly scanning and auditing on historical project datasets to identify potential data issues, cost overruns, or schedule anomalies for further manual investigation. - Environments where Python + pandas/numpy/scipy are already standard, and where minor customization (e.g., additional guards, threshold tuning, or integration of time-series detection into workflows) is acceptable.

Comments (0)

Post a Comment

No comments yet. Be the first!