AWS cost and reliability, for funded startups
AWS cost reductions your engineering team can act on.
You get an engineer-ready AWS savings backlog priced from your billing data. Your team can implement it directly, or I can ship approved items in a fixed-price sprint and verify the result against your bill.
How an engagement works
Assess
Read-only, ten business days. Evidence, risk, rollback guidance, and a verification method on every backlog item.
Implement, if needed
Your team can implement the backlog directly, or I can implement approved items in a separate fixed-price sprint and verify the result against billing data.
Monitor, optionally
Optional anomaly monitoring, monthly spend notes, and quarterly savings PRs.
Recommendations are not implementation plans
Cost Optimization Hub, Compute Optimizer, and related AWS services are useful inputs. The assessment validates their recommendations against your billing history and architecture, resolves overlaps, sequences dependent changes, and adds ownership, production risk, rollback, and post-change verification.
Without that context, valid recommendations compete with product work and stay unshipped. Engineers need to know the expected savings, the evidence behind them, the production risk, and the safest path to implementation.
The assessment turns billing and resource data into a ranked engineering backlog with those decisions already documented.
AWS Cost Assessment
From $5k fixed. Read-only access. Delivered as a written report within ten business days of access grant. The assessment leads with cost, includes reliability risks that affect safe implementation, and reviews AI workload costs where those workloads exist.
Spend baseline
A 13-month Cost Explorer trend, plus the most detailed current-quarter breakdown supported by your billing exports, tags, and allocation data. Any resolution limits are stated in the report.
Ranked savings backlog
Every material finding supported by the available data, priced in first-year dollars and scored for effort and production risk, with a rollback note and named verification method.
Commitment posture
Savings Plans and Reserved Instance coverage and utilization, upcoming expirations, and renewal decisions.
AI workload review
Where present, Bedrock, SageMaker, and GPU infrastructure costs, with workload or team attribution only where your existing tags and telemetry support it.
Reliability findings
Material exposures under growth or partial failure, each tied to available resource evidence and a business consequence.
Implementation options and walkthrough
Fixed-price implementation options mapped to your backlog, plus an included walkthrough and Q&A. The report is written for direct implementation by your engineering team.
Each recommendation includes the evidence needed to act
These three rows from the sample report show the expected savings, implementation effort, production risk, and verification method for each item.
| # | Finding | First-year | Effort | Risk |
|---|---|---|---|---|
| 1 | EKS right-sizing: workloads request 3.2x their P95 CPU usage; node utilization P95 34%. Fix requests, enable consolidation, conservative 25% node reduction. verify: EKS node line in CUR, staged by namespace, non-prod first | $46,000 | L | Medium |
| 2 | Non-prod scheduling: staging and dev compute runs at essentially the same level on weekends as on weekdays. Working-hours schedule, ~62% reduction. verify: daily-granularity spend on non-prod accounts; disable to revert | $24,000 | M | Low |
| 7 | Bedrock prompt caching on a 2.4k-token shared prefix. Modeled for GPT-5.6 Sol at a conservative 90% cache-hit rate. verify: cached-token share and cache-read spend after seven days | $6,000 | S | Low |
From the downloadable sample report: a synthetic organization at approximately $78k/month, built to demonstrate the method. $161k identified across 12 items. In a real engagement, each figure traces to a saved calculation over the available billing data.
From read-only access to a ranked backlog in ten business days
Scope and intake
We agree the scope and fixed price, complete the written intake and data-handling checklist, and confirm the read-only access plan.
Read-only access
An IAM role template is provided. The assessment uses billing and resource metadata only, with no permission to change production.
Analysis
I analyze Cost Explorer, available billing exports, commitments, resource utilization, and reliability signals. Every dollar figure comes from a saved billing-data calculation.
Report and walkthrough
You receive the ranked backlog, reliability findings, and implementation options, followed by a live walkthrough and Q&A. Access revocation is confirmed in writing.
Implementation is scoped directly from the backlog
The report is written for direct implementation by your team. If you want implementation support, each engagement has a fixed price agreed in writing before work begins and is limited to approved backlog items.
Assessment fee and implementation. Up to $5,000 of the assessment fee applies once to the first Cost Reduction or Reliability Hardening Sprint contracted within 90 days of report delivery. It does not apply to retainers.
-
$15k to $35k
Cost Reduction Sprint
Terraform PRs for approved backlog items, written approval before production changes, and a rollback plan for each change. Results are verified against billing data after the relevant reporting lag.
-
$10k to $25k
Reliability Hardening Sprint
Approved reliability findings such as failover remediation, alarm coverage on revenue paths, backup verification, and quota headroom. Runbooks and monitoring are included for ongoing operation.
-
from $2k/mo
Cost-Watch Retainer
Continuous billing monitoring with anomaly alerts, a monthly written spend note with deltas explained, and quarterly savings PRs as new opportunities appear.
-
from $3k/mo
Reliability Retainer
For systems hardened in a sprint: automated monitoring and escalation handling within an agreed service window and incident buffer. Sprint-first clients only.
Production infrastructure experience behind the assessment
modeled EKS versus ECS Fargate comparison
I modeled the company's production service fleet on its ECS Fargate configuration and on the EKS platform I had built. The model identified compute costs up to 80% lower on the workloads that benefited most.
regulated digital-asset trading platform, names withheld
- Multiple EKS platforms delivered and rolled out to production, including a platform for a third-party ledger system and a greenfield platform at a subsequent company.
- Datadog log-indexing spend reduced at the digital-asset trading company from the case study.
- More than seven years in fintech and digital-asset trading, including infrastructure supporting live order flow with strict availability requirements.
- Infrastructure as code throughout: Terraform and Pulumi delivery, OpenTelemetry and Grafana observability builds.
Who the assessment is for
This works when
- +You're funded or revenue-generating, roughly 30 to 300 people, with an established engineering team.
- +AWS spend is $15k/month or more. The sweet spot is $50k to $150k.
- +A technical sponsor owns the problem: CTO, VP Engineering, platform lead, or technical founder.
- +You'll grant read-only billing and resource access for the assessment.
Not a fit
- xUnder ~$10k/month AWS spend. Start with the free self-audit to identify areas worth investigating.
- xShopping for credits, resale discounts, or a managed billing layer.
- xRequiring guaranteed savings before implementation. Identified savings depend on which changes your team approves and ships.
- xA FinOps function already implementing and verifying recommendations regularly.
I do the assessment and implementation myself
I'm Justin Henderiks, a staff-level platform engineer with ten years in engineering and more than seven years in fintech and digital-asset trading. My work includes market data and order routing systems, AWS and Kubernetes platform architecture, and infrastructure defined in Pulumi and Terraform.
Engagements are async-first so scope, findings, and decisions stay documented. I'm available for calls whenever they help move the work forward, and every assessment includes a live walkthrough and Q&A after delivery.
Professional history: LinkedIn.