VeUP
← All case studies
Machine Learning Competency · FinServ AI/ML + FinOps
A Middle East financial-services AI firmIdentity protected

WarburgAI migrates a GPU-heavy FinServ AI workload to Amazon SageMaker

Cloud-to-cloud migrationManaged billing & resellGPU capacity & inference economicsCommitment & RI optimizationRightsizing & instance-family modernizationPer-service spend attributionStanding cost-optimization mechanism
~6 weeks
cloud-to-cloud migration, cutover included
Known number
GPU reserved via Capacity Blocks, inference on Graviton
Standing
managed billing and FinOps ever since
Amazon SageMakerCapacity Blocks for MLAWS GravitonAWS Cost Explorer

Shared anonymously — the customer’s name is held by VeUP and available on request.

GPU is the most expensive line on an AI company’s cloud bill — and the easiest to lose control of. VeUP moved this firm’s AI workload off another cloud and onto Amazon SageMaker, reserved its GPU through Capacity Blocks for ML, shifted inference to AWS Graviton, and wrapped the whole estate in a multi-account FinOps structure that made the spend visible, predictable, and governed.

The challenge

The firm — a Dubai-based, AI-driven financial-services company — ran its GPU/ML workload on a different cloud and wanted to scale on the AWS AI/ML ecosystem. But the move only made sense if the spend came with it under discipline. On-demand GPU is the most expensive way to run a bursty ML workload; a single-account billing posture meant nobody below root could see what anything cost; and a Marketplace-vs-direct spend mis-attribution had quietly distorted the whole cost picture.

The solution

VeUP ran the migration through all three phases — assess, mobilize, migrate — and landed the AI workload on Amazon SageMaker. Two cost decisions carry the design: GPU capacity for training and inference is reserved through Capacity Blocks for ML, trading unpredictable on-demand pricing for a known number, and inference runs on AWS Graviton for better price-performance. Underneath, a multi-account payer/child structure with Cost Explorer and a CUR-ingest cost platform gives the firm scoped, non-root cost visibility and a monthly optimization cadence — with VeUP running the engagement as managed billing and FinOps.

Production outcomes

KPIResult
Production outcomesThe AI workload runs in production on AWS. FinOps corrected the Marketplace-vs-direct spend mis-attribution that had distorted the cost picture, then shifted the compute mix toward Capacity Blocks for ML and Graviton inference — reducing the effective run cost of the workload, tracked continuously in CUR dashboards.
TimelineMigration kicked off in mid-December 2024; the workload was live on AWS by the end of January 2025 — about six weeks, cutover included. VeUP has run managed billing and FinOps for the firm since.
Cost postureCost engineering is the heart of the engagement: a monthly FinOps cadence works the levers — Capacity Blocks for ML over on-demand GPU, Graviton for inference — with a CUR-ingest cost platform and Cost Explorer providing visibility across the multi-account payer/child structure.
Lessons & continuationFor a GPU-heavy FinServ AI workload, the controlling cost levers are reserving GPU via Capacity Blocks for ML (not on-demand) and shifting inference to Graviton; a multi-account payer/child + CUR structure is the precondition for scoped, non-root cost discipline; correcting a Marketplace-vs-direct spend mis-attribution before optimizing avoids modeling against a distorted baseline.
AWS services in production
Amazon SageMakerCapacity Blocks for MLAWS GravitonAWS Cost ExplorerAWS Organizations (payer/child)AWS Cost & Usage Report

Architecture

Production AWS architecture: an AWS Organizations payer/child structure with CI/CD-driven Infrastructure as Code, a multi-AZ VPC running Amazon SageMaker training on Capacity Blocks for ML and AWS Graviton inference, ElastiCache and KMS-encrypted S3 for data and model artifacts, a security rail of IAM, KMS, CloudTrail, and security groups, and a CloudWatch, Cost Explorer, and CUR FinOps rail.
The architecture on AWS — SageMaker training on reserved Capacity Blocks, Graviton inference, and the FinOps rails that keep the GPU bill a known number.

Where it started

Assessed baseline · pre-migration source cloudFinancial-services AI · Middle East · workload on another cloud provider
Starting point
A FinServ AI workload off AWS

GPU/ML training and a live production model-serving loop running on the source cloud, with a Redis-style cache and object-backed data tier feeding the pipeline.

Gap
On-demand GPU only

No reserved capacity — the most expensive pattern for a bursty ML workload, with uncontrolled on-demand spend driving run cost.

Gap
A distorted cost baseline

Single-account billing with Marketplace-vs-direct mis-attribution — no scoped view of what the workload truly cost.

Constraint
No resilience on the serving path

A single serving loop with no multi-AZ posture for the live production model-serving path.

Driver
Locked out of the AWS AI/ML ecosystem

No access to Capacity Blocks for ML or Graviton economics while the workload stayed on the source cloud.

The source-cloud estate as assessed before the migration — the serving loop was later re-platformed onto Amazon SageMaker.