8.2
/ 10
1 evaluations
2.8k Downloads
Overview
Automatically detects cron jobs that failed due to transient network/connection errors and retries them on the next heartbeat, using the cron list/run tools and simple error-pattern matching.
Key Advantages
1.Hands-off recovery of failed cron jobs caused by network glitches, reducing manual intervention.
2.Integrates directly with HEARTBEAT.md workflow: minimal setup and clear implementation steps.
3.Uses explicit network error patterns to distinguish transient issues from logic/auth/config errors.
4.Respects cron job enabled state and avoids retrying disabled or recently-run jobs, reducing retry storms.
5.Provides both automatic (heartbeat) and manual recovery paths via cron CLI commands for debugging and control.
Use Cases
- Recovering message-delivery crons (e.g., Telegram, Slack, email digests) that fail due to temporary network outages.
- Ensuring periodic briefing/reporting jobs are delivered even when upstream APIs or network connections briefly drop.
- Automatically re-running external API polling or sync jobs that intermittently fail with ETIMEDOUT/ECONNREFUSED/ENOTFOUND.
- Maintaining reliability of notification or reminder systems where occasional DNS or socket errors are expected.
- Reducing operator toil in environments with unstable connectivity (e.g., remote deployments, flaky proxies) by auto-retrying transient failures.
Evaluation Scores
8.2
/ 10
Reliability
8.0
Functionality
8.2
Usability
8.8
Safety
7.8
Performance
8.8
Compatibility
8.0
Based on 1 evaluation · Latest: 3/19/2026
Download Trend
Loading...
Evaluation History (1)
8.2/103/19/2026▼
OS: darwin-x64LLM: moonshotai/kimi-k2.5
**Judgement**
A focused, well-scoped skill that adds robust auto-retry behavior for cron jobs that fail due to transient network issues. It integrates cleanly with the standard heartbeat + cron tooling model and should materially improve reliability of message- and API-driven jobs with minimal configuration.
**What it does well**
- Hooks into the heartbeat loop to scan `cron list` results and automatically re-run jobs with `lastStatus: "error"` when `lastError` matches known network/connection patterns (e.g., `ECONNREFUSED`, `ETIMEDOUT`, `Network request.*failed`, `sendMessage.*failed`, `fetch failed`, `socket hang up`).
- Clearly separates retryable transient errors (network, DNS, timeouts) from non-retryable ones (logic/config/auth issues, disabled jobs).
- Provides simple manual recovery commands (`cron list` + `cron run --id`) for targeted debugging.
- Mentions safeguards against retry loops (checks `lastRunAtMs`) and respects `enabled` state.
**Key risks / limitations**
- **False negatives:** Only errors matching the documented patterns get retried. Any non-standard or differently phrased network errors will be missed and require pattern updates.
- **False positives / side effects:** Some errors might be misclassified as "network-like"; rerunning the job could cause duplicate side effects for non-idempotent workflows (e.g., sending the same notification or performing a write twice).
- **Retry strategy is basic:** No configurable backoff, retry limits per job, or escalation; it simply retries on each heartbeat when conditions are met.
- **Assumes error-field consistency:** Relies on `job.state.lastError` being informative and consistent across tools/integrations.
**Recommended scenarios**
Use this skill when:
- Your cron jobs are primarily **idempotent** or safely repeatable (notifications, status checks, read-only syncs, report generation).
- You frequently see cron jobs fail with **transient network errors** (timeouts, DNS issues, connection resets) and want them retried automatically.
- You already use or are willing to adopt a **heartbeat-based health check** (via `HEARTBEAT.md`) and the standard `cron list`/`cron run` tools.
Be cautious or add extra guards when:
- Cron jobs perform **non-idempotent operations** (creating orders, charging cards, irreversible updates) where a second execution could be harmful.
- Error messages are highly variable or localized, making pattern matching less reliable.
Overall, this is a strong utility skill for improving reliability of network-dependent cron jobs in typical OpenClaw/heartbeat setups, with low operational overhead but some need for judgement around which jobs are safe to auto-retry.
Comments (0)
No comments yet. Be the first!