"We knew the cloud bill was too high, and we had the reports to prove it. What we didn't have was anyone actually turning those reports into a smaller invoice." - Client representative
A fast-growing US crypto wallet.
3 Alpacked Platform/DevOps engineers.
Since July 2026, ongoing.
The client is a fast-growing US crypto wallet used by 20 million people through a browser extension and mobile apps. The backend handles transaction submission, authentication, and integrations with multiple blockchains. Infrastructure spend at the company is discussed at board level.
Cloud and monitoring costs were growing faster than the business, and leadership set a target: $2M in annual savings. The company already had several reports on potential savings, but none of them had made the bill any smaller.
Three of our engineers work within the client's Backend Platform team, alongside the Developer Experience and Security teams. Cost optimization is an ongoing workstream with quarterly targets. This case covers work from July through September 26, 2026.
$948.7K/year in savings
0 user-facing incidents caused by cost optimization
$456K/year in August
< 1 week to replace a SaaS tool
Infrastructure costs for a platform handling around 200 million requests a day were growing faster than the business, and no one owned them. They had to be cut on a real-time financial platform with no maintenance windows, no development freeze, and not a single user-facing incident. The problem came down to three factors.
1
Challenge #1: Unused Resources and Uncapped Bills.
Some resources had outlived their purpose. RDS replicas were still running in two regions even though the services that needed them had long since moved to other regions. ElastiCache clusters sat idle, ECR was accumulating images with no lifecycle policy, and AWS Inspector was scanning accounts nobody used. Two third-party services billed with no upper limit: n8n, which handled PR and deployment notifications, and an EVM gas sponsorship provider. The only gas spend alert fired when the limit was reached, by which point the service had already stopped.
2
Challenge #2: Settings Nobody Revisited as the Platform Grew.
Datadog was collecting full container metrics, indexing info-level logs, and sampling traces at a rate chosen for a much smaller platform. Data in DynamoDB and S3 was stored without regard to how it was actually accessed.
3
Challenge #3: Costs Without an Owner.
Karpenter-provisioned nodes and shared services had no cost-allocation tags. That portion of spend couldn't be attributed to any team, so no one saw their own numbers or was accountable for them.
"We knew the cloud bill was too high, and we had the reports to prove it. What we didn't have was anyone actually turning those reports into a smaller invoice." - Client representative
This is a wallet holding users' funds, so availability and transaction correctness are the core of the product. If a transaction fails during a sharp market move, a person can't move their own money exactly when it matters most.
Load is uneven. A chain event, a product launch, or a partner integration can increase traffic by an order of magnitude within minutes. The infrastructure is correspondingly large:
Removing a replica or lowering sampling is technically easy. But a change like that can leave the on-call engineer without the signal they need during an incident, or remove capacity that will be needed at the next peak. So we worked by strict rules:
We ran cloud cost optimization as ongoing delivery work, not as an audit. Every saving got its own ticket in a dedicated project and was counted only after the reduction showed up on the bill. We measured results as annualized run-rate savings: how much costs would drop over a year if the current level holds. This made progress toward the target visible week by week and made it impossible to inflate the number or count the same saving twice.
First, we made costs visible. We extended tagging, including to Karpenter nodes, so teams could see most of the spend on their own resources. When a team has its own number, it also has accountability for it.
Then we tackled the two largest cost pools in parallel: unused AWS resources and Datadog bills. We treated Datadog as an engineering problem, not a procurement one: we looked for telemetry that gave teams no useful signal rather than for a vendor discount. We also reviewed storage separately. S3 data was moved into the right storage classes and cleaned up. For DynamoDB, it turned out that the access pattern for the terabytes of data stored there made cheaper storage worth considering.
In August, these workstreams delivered $456K/year:
Next, we addressed the uncapped bills. We either capped them with tiered alerting or replaced them with our own services.
The first step was rightsizing: smaller instances, tighter pod bin-packing, more aggressive autoscaling. It delivered results, but the main savings potential lay elsewhere: in observability bills and in stateful resources that no longer served their purpose. Rightsizing doesn't surface costs like these.
What we ruled out: reducing the fidelity of all telemetry at once. That approach would have produced bigger savings on paper but left on-call engineers without the signals they need. So we reviewed each type of telemetry separately, the same way we approach any log management and monitoring setup:
Lessons from the client's earlier attempt. The client had previously rolled out centralized database performance alerting across the company but abandoned it because there were too many alerts to act on. So we first tested slow-query monitoring on a single service using a monitor-plus-runbook pattern, then handed it to teams as a template each one configures with its own thresholds.
Each individual change is simple. The risk shows up later, during the next incident, when it turns out the signal you removed was exactly the one the on-call engineer needed. In August, the team closed around 68 tickets across several workstreams: cost, developer experience, access, database reliability, and ongoing platform support. That's why every cost-related change shipped separately, with its own justification and rollback path, rather than as one big release.
The second challenge was reporting credibility. Projected savings are easy to claim and hard to verify. Because we reconciled everything against the bills, the program reported a smaller number than it could have, but the finance team could verify every figure.
Banshee. An in-house notification service for PRs and deployments, built to replace n8n. We built it from scratch in under a week: repository, migration of the logic and data schema, Helm chart, CI, image registry with image-writer, GitOps, and ingress. Per-execution payments to n8n stopped. We deliberately left this saving out of the total because there was no confirmed dollar figure, even though the risk of an uncapped bill was gone.
GitHub webhook router refactor. While Banshee was in development, we moved the highest-volume workflows to a dedicated worker. n8n invocations dropped by about 85%, and the bill started going down before the switch to the new service.
Log volume notification service. It shows each team how many logs it generates and records whether the team acted on the notification. This turned tags into a feedback mechanism rather than just data for a report.
Spend limit alerts for the gas provider. A warning at 50% of the limit, a page to on-call at 90%, and repeat alerts while the condition stays active. Previously, the team learned the limit had been hit only after the service stopped. Now it gets warned before the limit runs out and has time to respond.
Access. During a security review of the mirrord rollout, we found a dangerous combination of permissions: cluster-admin plus cross-cluster access. It could have allowed unnoticed access to production data. The excess permissions were audited and removed.
Database. A brief cross-region loss of database connectivity was spotted right away by the on-call team. In under two minutes, they confirmed the connection had recovered on its own and all nodes were healthy. No incident was declared, and users noticed nothing.
CI. A CI outage blocked the shared build queue ahead of a cross-team branch cut. The team cleared over 100 queued jobs the same day, and the release that depended on them shipped on schedule.
Cloud & platform: AWS (EKS, RDS, Fargate, S3, ECR, ElastiCache, DynamoDB, Inspector, IAM, VPC), Kubernetes, Karpenter, Helm, ArgoCD / GitOps, Docker
Databases: CockroachDB, DynamoDB
Observability: Datadog (infrastructure, APM, logs)
IaC & CI/CD: Terraform, GitHub Actions
Other: mirrord, Linear
Cost optimization had no effect on how the wallet worked: the platform continued to serve 20 million users and around 200 million requests a day. What changed was elsewhere: the bill started shrinking, and teams got visibility into what their services cost. Here are the key results:
1
Cost Reduction
As of September 26, 2026, annualized run-rate savings stand at about $948.7K, nearly half of the $2M target. August alone added $456K/year: $186K on AWS and $270K on Datadog. Every figure was verified against the bills and tied to a closed ticket. After the move to Banshee, the company no longer pays n8n per execution. This saving isn't included in the total because there is no confirmed dollar figure. With extended tagging, teams can see most of the spend on their services and manage it.
2
Delivery Speed
In August, the team closed around 68 tickets across several workstreams while also supporting the platform. Banshee went from architecture investigation to production with CI, GitOps, and ingress in under a week. After a CI outage, over 100 jobs in the shared queue were cleared the same day, and the release that depended on them shipped on schedule.
3
Reliability & Security
Cost optimization caused no user-facing incidents. Excess cluster-admin permissions that could have allowed unnoticed access to production data were removed. The team now gets spend-limit warnings for the gas provider at 50% and 90% instead of after the service stops.
We'll help you find the main sources of overspending and cut costs while keeping production risk under control.
Just fill the form below and we will contaсt you via email to arrange a free call to discuss your project and estimates.