Stuck data jobs do not throw errors. They just bill you. ZombieJob watches CloudWatch, spots the ones that died quietly, and hands you a kill switch in Slack.
No credit card. One AWS account free, forever.
🧟 Zombie job detected
CPU at 0.4% for 30 minutes. No log output for 94 minutes.
CloudWatch alarms fire on errors. A hung Spark executor never throws one.
Three rules, no black box. Runtime past 3x baseline, CPU under 5 percent, or logs silent for 30 minutes.
Every alert carries the DPU-hours burned since the job should have finished. Priced at current AWS list rates.
Alerts land with the job name, the overrun, the spend, and a kill button. No dashboard hunting at 2am.
Run one CloudFormation stack. We hold a role ARN and an external ID, assumed at runtime with STS.
Median of the last 30 successful runs, refreshed nightly. Outliers stop poisoning the threshold.
A stuck Glue job gets flagged inside 10 minutes of crossing the line. Not the next morning.
Run our CloudFormation stack. Read-only plus a stop permission.
Discovery pulls every Glue job and builds baselines from run history.
Detection runs every 5 minutes against live CloudWatch metrics.
One click from Slack or the dashboard. The abort is logged and audited.
One stuck job pays for a year of Pro.
Prove it works on one account.
For teams running production pipelines.
$1,490 a year, two months free
Shared ownership across a data org.
Up to 5 seats
Connect an account and see your first baseline in under ten minutes.
Start free