Comparison
Datadog vs Grafana: what each really costs, and the third option.
Pricing model, what drives the bill, operations effort and migration cost, compared without a vendor in the room. Then the option most teams skip: fix the Datadog bill first and decide with a smaller number.
A free 30-minute demo with an engineer. No commitment.


At a glance
Full comparison below. No unit prices for Grafana or New Relic are quoted here; their models change and the shape of your estate matters more.
Side by side
Three ways to pay for observability.
Datadog list prices are the ones our calculator uses: $15 to $23 per infrastructure host, $31 per APM host, $0.05 per custom metric beyond the allotment, $0.10 per gigabyte ingested and $1.70 per million events indexed at 15-day retention, all on annual terms.
| Datadog | Grafana Cloud | Self-hosted Grafana stack | |
|---|---|---|---|
| Pricing model | Per host for infrastructure and APM, per custom metric beyond the allotment, per gigabyte ingested, per million events indexed. Annual commitments with on-demand overage. | Usage-based: active metric series, log volume, trace volume and users, with a free tier and a paid tier above it. | Open-source Grafana with Loki, Tempo and Mimir or Prometheus. No licence. You pay for compute, storage and the people who run it. |
| What drives the bill | Host counts, custom metric cardinality, log indexing volume and retention. Defaults favour the meter. | Series cardinality still counts, plus log and trace gigabytes and retention. The same tag problem follows you. | Storage and compute for the volume you keep, and engineering time. Cardinality still hurts, as memory rather than invoice. |
| Setup and operations effort | Install the agent and it works. Integrations, dashboards and monitors come ready. Little to operate. | Managed backend. You still build dashboards, alerting and integrations, and instrument with Grafana Alloy or OpenTelemetry. | Highest. Cluster sizing, upgrades, retention, backups and on-call for the observability stack itself. |
| Logs, metrics and traces | All three in one product with correlation built in, plus RUM, security and synthetics as add-ons. | All three through Loki, Mimir and Tempo, correlated in Grafana. Coverage depends on what you instrument. | The same components, at the version and scale you choose to run. |
| Alerting and on-call | Monitors, SLOs and on-call in the product. | Grafana Alerting and Grafana OnCall, or an external pager. | Grafana Alerting plus whatever pager you wire in. Nothing is on by default. |
| Who it suits | Teams that want one vendor and would rather pay than operate, if the usage is kept in check. | Teams with Prometheus habits, moderate cardinality and someone who owns dashboards. | Teams with a platform group that already runs stateful systems and wants control over data and cost. |
| Migration cost | Staying costs nothing to move. Leaving means rebuilding everything above. | Dashboards, monitors, SLOs, integrations and runbooks rebuilt. Months of parallel running. | The same rebuild, plus standing up and hardening the stack before the parallel period starts. |
Where it hurts
Where teams get burned on each.
Datadog
Cardinality
A high-cardinality tag on one metric becomes thousands of billable series. The invoice moves before anyone notices the tag.
Indexing
Ingesting is cheap, indexing is not. Indexing every line is the default, so the log bill tracks traffic rather than what anyone searches.
Host sprawl
Agents on every machine, APM on hosts that never needed tracing, and committed counts that outlive the estate they were sized for.
Grafana
Series cardinality still costs
Grafana Cloud bills on active series. The tag that broke the Datadog bill breaks this one too, and self-hosted Mimir pays for it in memory.
Log queries and retention
Loki is cheap to store and expensive to query badly. Retention decisions and label design set the bill and the query speed.
On call for your own observability
Self-hosting means upgrades, storage, backups and a pager for the system that is supposed to tell you when things break.
The move
What a migration actually costs.
The licence is the small number. The engineering time and the overlap are the large ones.
- Every dashboard, monitor and SLO rebuilt by hand or scripted, then verified against the old one.
- Integrations and agents replaced, from cloud integrations to the tracing libraries in your services.
- Runbooks rewritten and the on-call rotation retrained on new tools before they are paged on them.
- Months of running both stacks in parallel, with two sets of alerts to keep honest.
- The second bill during the overlap, and the old commitment if the renewal has already passed.
The third option
Cut the Datadog bill first.
Most Datadog bills carry usage nobody decided on: cardinality, indexing defaults, agents on hosts that never needed them, and commitments sized for last year. A free, read-only audit finds it in seven days, and in production accounts we have run, the cut reaches up to 40%.
The audit sets the number to beat
A migration has to beat Datadog with the defaults fixed, not Datadog as it is today. Most teams find the fixed bill changes the decision.
How the Datadog cost audit worksIf you do switch, we run it
Iris works across Datadog, CloudWatch and Grafana. The migration is a scoped project, the parallel period runs under one set of alerts, and the new stack is operated on a fixed monthly fee.
See the managed serviceQuestions
Datadog, Grafana and the move between them.
Is Grafana cheaper than Datadog?
It depends on what drives your bill. Grafana Cloud is usage-based rather than per host, which helps estates with many small hosts. It still bills on active series and on log and trace volume, so a cardinality problem follows you across. Self-hosting removes the licence and adds engineering time. The audit tells you which cost you actually have.
Can we run both?
Yes, and most migrations do for a while. A common end state is Datadog for APM and a Grafana stack for metrics and logs, or the reverse. Iris reads across Datadog, CloudWatch and Grafana, so the split can be run under one set of alerts.
What about New Relic?
The same conversation with a different pricing model, usage-based on data ingested and on users. The questions are identical: what drives the bill, what the migration costs, and whether cutting the current bill first changes the answer.
How long does a migration take?
It depends on how many dashboards, monitors and integrations you have and how much of it is in code. Plan for a scoped project followed by months of parallel running rather than a switch-over weekend.
What do you recommend?
Run the free audit first. It shows what your Datadog bill would be with the defaults fixed, which is the number a migration has to beat. If a move still wins, we scope it with the audit in hand and run the parallel period.
Decide with a smaller number.
Book a demo with an engineer. Bring last month’s Datadog invoice and you’ll get a first read on what the defaults are costing you.
Estimate first with the Datadog pricing calculator.
You’ll pick a time on the next page.