Now monitoring AWS Glue, EMR and Lambda

Your Glue job has been
RUNNING for 6 hours.
Nobody noticed.

Stuck data jobs do not throw errors. They just bill you. ZombieJob watches CloudWatch, spots the ones that died quietly, and hands you a kill switch in Slack.

No credit card. One AWS account free, forever.

#data-alerts

🧟 Zombie job detected

Jobcustomer-etl-nightlyRuntime6h 12mBaseline41mWasted so far$47.30

CPU at 0.4% for 30 minutes. No log output for 94 minutes.

Abort jobView details

Built for the jobs that fail without failing

CloudWatch alarms fire on errors. A hung Spark executor never throws one.

Deterministic detection

Three rules, no black box. Runtime past 3x baseline, CPU under 5 percent, or logs silent for 30 minutes.

Real dollar figures

Every alert carries the DPU-hours burned since the job should have finished. Priced at current AWS list rates.

Slack that does something

Alerts land with the job name, the overrun, the spend, and a kill button. No dashboard hunting at 2am.

No keys stored

Run one CloudFormation stack. We hold a role ARN and an external ID, assumed at runtime with STS.

Baselines that learn

Median of the last 30 successful runs, refreshed nightly. Outliers stop poisoning the threshold.

Five minute polling

A stuck Glue job gets flagged inside 10 minutes of crossing the line. Not the next morning.

Four steps, one afternoon

01

Connect an account

Run our CloudFormation stack. Read-only plus a stop permission.

02

We learn your jobs

Discovery pulls every Glue job and builds baselines from run history.

03

Zombies get flagged

Detection runs every 5 minutes against live CloudWatch metrics.

04

You kill it

One click from Slack or the dashboard. The abort is logged and audited.

Pricing

One stuck job pays for a year of Pro.

Free

Prove it works on one account.

$0forever

  • 1 AWS connection
  • 10 monitored jobs
  • Real-time detection every 5 minutes
  • Slack alerts with wasted cost
  • 7 day baseline history
Start free
Recommended

Pro

For teams running production pipelines.

$149per month

$1,490 a year, two months free

  • Unlimited AWS connections
  • Unlimited monitored jobs
  • One-click abort from Slack
  • 30 day baseline history
  • Cost savings dashboard
Upgrade to Pro

Team

Shared ownership across a data org.

$399per month

Up to 5 seats

  • Everything in Pro
  • Up to 5 team members
  • Custom detection rules
  • Audit log
  • SSO
  • Priority support
Start with Team

Find out what is running right now

Connect an account and see your first baseline in under ten minutes.

Start free