ClawTrust LogoClawTrust
Cron Retry

Cron Retry

by jrbobbyhansen-pixel · v1.0.0

Design
ClawHub
8.2
/ 10
1 evaluations
2.8k Downloads

Overview

Automatically detects cron jobs that failed due to transient network/connection errors and retries them on the next heartbeat, using the cron list/run tools and simple error-pattern matching.

Key Advantages

1.Hands-off recovery of failed cron jobs caused by network glitches, reducing manual intervention.
2.Integrates directly with HEARTBEAT.md workflow: minimal setup and clear implementation steps.
3.Uses explicit network error patterns to distinguish transient issues from logic/auth/config errors.
4.Respects cron job enabled state and avoids retrying disabled or recently-run jobs, reducing retry storms.
5.Provides both automatic (heartbeat) and manual recovery paths via cron CLI commands for debugging and control.

Use Cases

  • Recovering message-delivery crons (e.g., Telegram, Slack, email digests) that fail due to temporary network outages.
  • Ensuring periodic briefing/reporting jobs are delivered even when upstream APIs or network connections briefly drop.
  • Automatically re-running external API polling or sync jobs that intermittently fail with ETIMEDOUT/ECONNREFUSED/ENOTFOUND.
  • Maintaining reliability of notification or reminder systems where occasional DNS or socket errors are expected.
  • Reducing operator toil in environments with unstable connectivity (e.g., remote deployments, flaky proxies) by auto-retrying transient failures.

Evaluation Scores

8.2
/ 10
Reliability
8.0
Functionality
8.2
Usability
8.8
Safety
7.8
Performance
8.8
Compatibility
8.0

Based on 1 evaluation · Latest: 3/19/2026

Download Trend

Loading...

Evaluation History (1)

8.2/103/19/2026
▼
OS: darwin-x64LLM: moonshotai/kimi-k2.5
**Judgement** A focused, well-scoped skill that adds robust auto-retry behavior for cron jobs that fail due to transient network issues. It integrates cleanly with the standard heartbeat + cron tooling model and should materially improve reliability of message- and API-driven jobs with minimal configuration. **What it does well** - Hooks into the heartbeat loop to scan `cron list` results and automatically re-run jobs with `lastStatus: "error"` when `lastError` matches known network/connection patterns (e.g., `ECONNREFUSED`, `ETIMEDOUT`, `Network request.*failed`, `sendMessage.*failed`, `fetch failed`, `socket hang up`). - Clearly separates retryable transient errors (network, DNS, timeouts) from non-retryable ones (logic/config/auth issues, disabled jobs). - Provides simple manual recovery commands (`cron list` + `cron run --id`) for targeted debugging. - Mentions safeguards against retry loops (checks `lastRunAtMs`) and respects `enabled` state. **Key risks / limitations** - **False negatives:** Only errors matching the documented patterns get retried. Any non-standard or differently phrased network errors will be missed and require pattern updates. - **False positives / side effects:** Some errors might be misclassified as "network-like"; rerunning the job could cause duplicate side effects for non-idempotent workflows (e.g., sending the same notification or performing a write twice). - **Retry strategy is basic:** No configurable backoff, retry limits per job, or escalation; it simply retries on each heartbeat when conditions are met. - **Assumes error-field consistency:** Relies on `job.state.lastError` being informative and consistent across tools/integrations. **Recommended scenarios** Use this skill when: - Your cron jobs are primarily **idempotent** or safely repeatable (notifications, status checks, read-only syncs, report generation). - You frequently see cron jobs fail with **transient network errors** (timeouts, DNS issues, connection resets) and want them retried automatically. - You already use or are willing to adopt a **heartbeat-based health check** (via `HEARTBEAT.md`) and the standard `cron list`/`cron run` tools. Be cautious or add extra guards when: - Cron jobs perform **non-idempotent operations** (creating orders, charging cards, irreversible updates) where a second execution could be harmful. - Error messages are highly variable or localized, making pattern matching less reliable. Overall, this is a strong utility skill for improving reliability of network-dependent cron jobs in typical OpenClaw/heartbeat setups, with low operational overhead but some need for judgement around which jobs are safe to auto-retry.

Comments (0)

Post a Comment

No comments yet. Be the first!