Managed Kubernetes
Your Kubernetes, run around the clock. One fixed monthly fee.
EKS or self-managed, upgraded, patched, scaled and watched by four agents, with senior engineers approving every change and holding the pager under a 99.9% uptime SLA.


- Sage AI
Cluster version leaves standard support in 41 days. Upgrade pull request opened: add-ons first, then node groups.
Scheduled for the Tuesday change window - Iris AI
checkout-api restarting, OOMKilled after the 23:30 deploy raised the heap without the limit. Paged on-call.
Resolved 01:05. Limit corrected in a pull request. - Finly AI
Requests on 14 workloads set well above usage. Rightsizing pull request opened, two nodes fewer at peak.
Awaiting approval - ClearRisk
CVE in a base image affects 2 running images. Patch pull request opened.
Approved by the on-call engineer, 03:40 - Summary
0 open incidents. 30-day uptime 99.97%. 3 changes, 2 approvals, 0 rollbacks.
Posted to #ops
What we run
Everything the cluster needs, kept in code.
Managed Kubernetes as one service: the control plane, the nodes, the platform add-ons, the identity and the bill, with one team accountable for all of it.
Upgrades and patching
Control plane and node upgrades on the EKS release cadence, before extended support fees apply. Managed add-ons kept current, node images rotated, and every version pinned in Terraform.
Sage AINode groups and autoscaling
Node groups sized to measured usage, with Karpenter or the cluster autoscaler tuned so capacity follows load and scales back down at night.
Finly AIIngress, DNS and certificates
The load balancer controller, external DNS and cert-manager wired in once and kept in code, so a new service gets a hostname and a certificate from a pull request.
Sage AIIdentity, secrets and network policy
IAM roles for service accounts or Pod Identity per workload, secrets pulled from your secret store rather than pasted into manifests, and network policy between namespaces.
ClearRiskObservability and on-call
Iris tunes the alerts so a page means customer impact, not a restarting pod. Engineers hold the pager under the SLA and every incident ends with a fix in code.
Iris AICluster cost
Requests and limits set from real usage, workloads packed onto fewer nodes, Spot where the workload allows it, and a monthly reconciliation against your invoice.
Finly AIHow a change happens
Agents handle the routine. Engineers own the outcome.
Nothing changes on a cluster without a named engineer approving it. Here is who does what when the usual things happen.
| When this happens | The agent | The engineer |
|---|---|---|
| A version reaches end of standard support | Sage AI Opens the upgrade pull request: control plane, add-ons and node groups in order, with the deprecation check attached. | Reviews the plan. Staging first, then production in the change window. |
| A node group runs hot | Finly AI Proposes new instance sizes and autoscaler bounds from 30 days of utilization, with the cost and the risk. | Approves. The change lands one node group at a time. |
| A pod is OOMKilled at 3am | Iris AI Correlates the restarts with the deploy that changed the memory limit and pages on-call with the diff. | Responds within the SLA, fixes the limit in code, writes the review. |
| A CVE lands in a base image | ClearRisk Checks which running images carry it and opens a patch pull request for the ones that do. | Merges the patch. The rest is filed, not paged. |
| Someone kubectl-edits production | Sage AI Detects the drift from the state in Git, reverts it and posts the diff. | Confirms, or turns it into a real change. |
| An auditor asks who can exec into pods | Sage AI Exports the RBAC bindings, the access reviews and the audit log for the period. | Answers the auditor with the export. |
EKS specifics
The parts of EKS that bite when nobody owns them.
Most cluster incidents we inherit trace back to a version nobody scheduled, an add-on nobody upgraded, or a change nobody wrote down. These are handled on a calendar, not when they break.
- Version upgrades scheduled before extended support fees apply
- Managed add-on upgrades: VPC CNI, CoreDNS, kube-proxy, the EBS CSI driver
- Control plane logging to CloudWatch, with retention set on purpose
- A cluster per environment or namespaces per team, decided with you and written down
- GitOps with Argo CD or Flux where you have it, pull requests where you do not
- Every cluster, node group and add-on in Terraform, so nothing lives only in a console
Onboarding
Live in four weeks.
A read-only review first, then reconciliation and hardening, then handover. Your team keeps shipping throughout.
- Week 1
Review and scorecard
A read-only kubeconfig and a cross-account role. We review versions, add-ons, RBAC, autoscaling, cost and alerting, and hand you a scorecard with what we would fix first.
- Weeks 2 to 3
Reconcile and harden
Cluster state reconciled into Terraform and Git, upgrades scheduled, node groups resized, alerts tuned, runbooks written, and our engineers join your rotation.
- Week 4
Handover to 24/7
Agents live under policy, engineers take the pager, the SLA starts. Your team goes back to shipping.
- Ongoing
Run and improve
Upgrades on the release cadence, monthly cost reconciliation against invoices, audit evidence always current, and a roadmap for what changes next.
Pricing
One fixed monthly fee, scoped to your clusters.
Kubernetes operations sit in the Growth tier and up. The fee is quoted after the free review, covers all four agents and senior engineer approval on every change, and runs month to month with no hourly billing.
Starter
- All four agents under policy
- Terraform ownership and drift control
- CI/CD with security gates
- 24/7 incident response, 99.9% uptime SLA
- Monthly cost reconciliation
Growth
- Everything in Starter
- EKS operations and upgrades
- Compliance automation with evidence export
- SOC 2 readiness in six weeks
- Disaster recovery, tested on a schedule
Scale
- Everything in Growth
- A named lead engineer
- Custom response and resolution targets
- Migration and re-platforming projects
- Quarterly architecture reviews
Only need the cluster bill fixed? The AWS cost audit is free and read-only and covers Kubernetes requests, limits and node groups.
From customers
What engineering leaders say.
“Since bringing DevLift's AI agents into our production environments, our infrastructure operates seamlessly. We established strict, automated security guardrails without slowing down our deployments, and Finly optimized our cloud waste by thousands of dollars automatically.”
“DevLift completely removed the ops toil from our sprint cycles. Instead of manually wrestling with IaC and chasing compliance drift, we rely on their platform to keep our environments stable and audit-ready. It's like having a senior SRE on staff 24/7.”
Case studies
What running it looks like.
Three engagements: the infrastructure chores that went away, and what the invoices and the auditors saw next.
Security guardrails on every deploy, and cloud waste cut automatically.
Read the case study
Infrastructure chores out of the sprint, and environments that stay audit-ready.
Read the case study
Regulated money movement, without paying for two of everything.
Read the case studyTrack record
Built by operators, not theorists.
We ran production clusters for fintech and crypto companies for four years before we encoded the recurring work into agents. The engineers who did it are the ones on your rotation.
The people behind the agents
DevLift is built by a small DevSecOps team in Dubai. Everyone touches production, everyone talks to clients, and the engineers who operate your accounts are the same ones who answer the page.
- Headquartered in Dubai. Runs production for fintech, crypto and AI companies across the UAE, the US and the UK.
- OSWE-certified engineers, DEF CON and Black Hat speakers and former CTF leads, with five discovered CVEs between them.
- Four years operating infrastructure by hand before the recurring work was encoded into the agents.
- A shared pager. Someone is reachable when it breaks, under an SLA, and the same engineers approve every agent change.


The DevLift team, 2026.
Founder
Harshil Olavakott
Founder and CEO, DevLift. LinkedIn
Cloud architect and DevSecOps specialist, building and running infrastructure on AWS and Azure since 2012. He has led engineering at Dentsu and ran production for fintech and crypto teams for four years before turning that work into DevLift.
Has built and run infrastructure for

Questions
Before you book.
Is it EKS only?
Mostly EKS, because that is where most of our customers run. Sage is multi-cloud, so GKE, AKS and self-managed clusters are covered by the same agents and the same rotation.
Do we keep kubectl access?
Yes. It is your cluster in your account. Your engineers keep their access, and drift detection tells us when a console or kubectl change needs to become a real change in code.
Can you migrate us onto Kubernetes?
Yes, as a scoped project first: from EC2, ECS or a legacy setup onto EKS, planned in the review and landed in Terraform. The managed plan then picks it up on day one.
Who holds the pager?
DevLift engineers in Dubai, on call around the clock for teams in the UAE, the US and the UK, under a 99.9% uptime SLA. Response and resolution targets by severity are written into the agreement before we start.
What about cluster cost?
Finly does the same work it does on the rest of your AWS: requests and limits from real usage, node groups sized and scaled, Spot where the workload allows it. The cost audit is built into the plan and reconciled monthly against your invoice.
What happens if we leave?
The plan runs month to month. The cluster is in your account and its state is in your repositories, so leaving means removing our role and taking the pager back.
See your clusters run this way.
Book a free 30-minute demo with an engineer. We'll show the agents on a live cluster and agree what a read-only review of yours would cover.

