Kubernetes consulting.
Set up, taken over, run.
Senior Kubernetes work at every stage: new clusters designed and hardened, inherited clusters mapped and stabilized, production operated with observability you can actually trust. And when you don’t need Kubernetes — we say so.
Clusters outlive the people who built them
Kubernetes usually arrives in a growth phase — set up quickly, extended under pressure, documented in nobody’s spare time. It works, so it ships. Then the engineers who understood it move on.
What remains is a cluster the team is afraid to touch: upgrades postponed past end-of-life, alerts nobody trusts, logs pulled from pods by hand, and a quiet hope that nothing fails on a Friday.
That is the state we take clusters over in — and the state we design new ones to never reach: as code, documented, observable, and upgradable by whoever comes next.
Where we start
Most engagements begin in one of three places — the path converges on the same steady state.
A new platform
EKS designed and built right-sized to your stage — as Terraform, with hardening, observability, and a clear path from commit to cluster from day one.
A takeover
An inherited or undocumented cluster mapped, stabilized, and documented — read-only first, before anything changes.
Steady-state operations
We run the platform: upgrades, capacity, alerts, and response — fractional days per month or fully managed.
What we handle
The standing Kubernetes work — from the control plane to the pager.
Cluster setup & EKS
Clusters provisioned as code on EKS — with GKE, AKS, and self-managed where you already run.
Security hardening
RBAC that means something, network policies, pod security standards, and IRSA — reviewed and enforced.
GitOps & deploys
Deployments from commit to cluster with reviewable manifests and rollbacks you can trust.
Observability — LGTM
Loki, Grafana, Tempo, and Mimir wired to every cluster, node, and pod — dashboards and alerts your team believes.
Upgrades without fear
Version upgrades staged, tested against your workloads, and reversible — no more clusters stuck at end-of-life.
Cost & autoscaling
Requests and limits from real usage, node autoscaling, and spot where it is safe — Kubernetes without the surprise bill.
Everything as code
Cluster, add-ons, and workloads in Terraform and versioned manifests — reproducible from an empty account.
Stateful workloads
Databases, queues, and storage run on or beside the cluster deliberately — not by accident.
Stabilize first. Then improve.
The same order whether the cluster is new to us or new to everyone.
Audit
Workloads, versions, RBAC, and single points of failure mapped — a full picture of what you actually run.
Stabilize
Observability, backups, and runbooks first — no changes until failures would be visible and recoverable.
Harden
Access, network policies, and the upgrade path fixed — the risky work, done deliberately.
Optimize
Requests and limits right-sized, autoscaling tuned, and cost mapped to workloads.
Run
Ongoing operations at the level you choose — fractional days or fully managed.
The team that knows how it runs
“We bought a product without the people who built it. Lodemark became the team that knows how it runs.”
Frequently asked questions
Do we actually need Kubernetes?
Not always — and we will tell you when you don’t. Below a certain scale, simpler platforms cost less and fail less. When you do need it, we build it so it stays operable after we step back.
Can you take over a cluster with no documentation?
Yes — that is the most common takeover. We map it read-only first: workloads, configs, RBAC, and data flows. Nothing changes until the audit is complete and failures would be visible.
Which clouds and distributions do you cover?
EKS is the deepest expertise; we also run GKE, AKS, and self-managed clusters — and migrate between them.
How do upgrades work without downtime?
Staged: the new version is tested against your workloads first, then nodes roll gradually with health checks and a rollback path. Clusters stuck past end-of-life get a catch-up plan, one version at a time.
What observability stack do you use?
Grafana’s LGTM stack — Loki for logs, Grafana for dashboards, Tempo for traces, Mimir for metrics — Prometheus-compatible and portable. If you already run Datadog or similar, we work with it.
Who owns the cluster and the code?
You do — always. Everything lives in your cloud accounts and repositories as code, documented so your next engineer can run it.
Beyond the cluster
Scale without chaos.
Tell us about your infrastructure - we'll reply within one business day.
Book a 30-minute intro call- 01A 30-minute intro call — your stack, your goals, no pitch deck.
- 02Read-only access (NDA first if you prefer) and about a week of analysis.
- 03A written report of findings and what we would fix first — yours to keep, either way.