Skip to content
Refine

Product · AWS Advanced Tier Services Partner

Catch unusual AWS spend before it compounds

Statistical baselines per service plus an AI commentary layer that explains why the line moved. Free, multi-account, multi-region.

How It Works

Statistical Baseline + AI Commentary

Fewer false positives than naive threshold alerts. Plain-English explanations on every alert so on-call does not have to dig.

  1. 1

    Establish Baseline

    Rolling 30-day window per service per account, calibrated to ignore weekday-vs-weekend and end-of-month invoice patterns.

  2. 2

    Detect Deviations

    Robust statistical thresholds (median absolute deviation, not naive mean) so noisy workloads do not bury real signals.

  3. 3

    Apply AI Commentary

    A natural-language layer turns the raw anomaly into a one-sentence explanation: "spike in running instances, possible missed shutdown scheduling."

  4. 4

    Route To Recipients

    Alerts route to the email addresses you configure per account, and to Slack or Microsoft Teams channels via incoming webhooks.

Refine's Anomaly Detection page: severity counters, two open new-spend anomalies on Amazon EC2 with the instance named, and a plain-English explanation of each
Near-real-time detection from a 7-day rolling z-score baseline. Each anomaly names the service and resource, the size of the move, and explains in plain English what usually causes it.

Examples from the product

What an Anomaly looks like in your Inbox

Each card is the exact shape of the notification you receive — service + region + delta + AI commentary.

  • Amazon EC2

    us-east-1Medium

    +7% week-over-week

    Spike in running instances starts Friday 19:00 and continues through weekend. Possible missed shutdown scheduling on the staging cluster.

  • AWS Lambda

    us-east-1High

    +152% in 24 hours

    Function `nightly-report-runner` invoked at 2.5× normal volume. Likely a runaway loop — concurrency limit not set. Worth checking the invocation source.

  • Amazon S3

    eu-central-1Medium

    +23% on storage

    Standard-tier growth on logs-prod bucket. Lifecycle policy may have regressed — last successful Glacier transition was 14 days ago.

  • Data Transfer

    cross-regionHigh

    +34% cross-region

    New traffic flow us-east-1 → ap-south-1 starting Tuesday. Check VPC peering, CloudFront origin, or recent deployment touching cross-region replication.

  • Amazon RDS

    us-east-1Info

    -100% (likely deletion)

    Instance `analytics-prod-replica` disappeared from the bill. If intentional, dismiss. If not, this is a costly accident — check CloudTrail before backups expire.

Notification channels

Email, Slack, and Microsoft Teams. All Live.

Three channels, all working today — no "integrations page" full of logos that turn out to be roadmap items.

Live

Email

Multi-recipient routing. Per-account or org-wide. Severity threshold per recipient.

Live

Slack

Incoming-webhook posts (Block Kit) with per-alert-type routing to the channel you choose.

Live

Microsoft Teams

Incoming-webhook card alerts that link straight to the anomaly in the dashboard.

Deciding what goes where

Growth and up

Three channels is the easy part. The reason alerts get muted wholesale at most companies is that everything goes everywhere, so routing is per event type, not per product.

  • A matrix of event type × channel, so a cost anomaly and a critical security finding need not land in the same place
  • Quiet hours, so an alert that can wait until morning does
  • Per-event mutes, for the anomaly you have already actioned and do not need reminding about
  • Webhook URLs encrypted with KMS in your browser, decrypted only inside the dispatcher
  • Editing a rule never re-sends an unchanged webhook — the plaintext does not round-trip
  • A webhook that exhausts its retries shows up for an admin, not in a log nobody reads

Tune sensitivity per service

Some workloads are inherently noisy — batch jobs, ML training, dev environments. Refine lets you tune thresholds per service so the noise stays quiet and the real signal cuts through.

  • Per-service sensitivity (Low / Medium / High / Off)
  • Preferences per account, so non-production can be quieter than prod
  • Quiet hours, so an alert that can wait until morning does
  • Mute a whole event type, or just one service or budget
  • Mutes can expire on their own, so a temporary silence does not become permanent
  • Muted alerts stay visible in-app — they stop leaving the building, not existing

Frequently Asked Questions

  • Refine builds a 30-day rolling baseline once you connect. New accounts see initial low-confidence alerts within 7 days; full-confidence anomalies after 30 days of history. You can also feed historical CUR to bootstrap immediately.

Stop finding cost spikes on the invoice

60-second setup. Free under $2,000/month of AWS spend. Read-only access.

Refine is built and supported by HabileLabs, an AWS Advanced Tier Services Partner.