Every engagement lands on Bedrock. The model has to earn it.
Claude, GPT, open weights and first-party Nova all arrive through one governed channel — and which one ships is settled by benchmark on the customer's own labelled data, workload by workload. Pick a provider to see the proof behind it.
Claude ships as the reasoning tier, the agent brain, and the qualified fallback — and runs VeUP's own delivery engine.
Open the Anthropic Claude practice- Published case studies
- 7
- Accounts in the estate
- 35
- Linked AWS accounts
- 168
published case studies
accounts in the work
linked AWS accounts
Published proof — Anthropic Claude
All 7 case studies →Running in productionThree agents, three models. Claude 3.5 Sonnet carries the reasoning steps and Claude Haiku the latency-sensitive ones — a split settled by benchmark against Amazon Nova before launch, then hardened with Bedrock Guardrails.Read the case study →
Running in productionClaude on Bedrock behind MCP tool servers over the CRM's existing REST APIs, with per-tenant context isolation and a human in the loop on v1.Read the case study →
Production — premium tierClaude Sonnet is the top rung of a three-step ladder — Nova Lite triages, Nova Pro answers by default, Claude Sonnet takes the questions that earn it. Answers land in 1–2s against a 5s SLA.Read the case study →
A published case study is the only production claim on this site — named services, measured outcomes, acceptance criteria. Everything else on these pages is the estate around it: accounts where a model family is under evaluation or in the conversation. Under evaluation is not in production, and this site never blurs the two.
Five candidates. One labelled dataset. The open model won.
Multimodal content moderation at platform scale for a global consumer social platform. This is what model-agnostic looks like when it is real: VeUP's own partner model was in the run and did not win it.
- SelectedQwen 3 VLOpen-weight
- EvaluatedMeta Llama 4 MaverickOpen-weight
- EvaluatedMeta Llama Guard 4Open-weight
- EvaluatedClaude OpusAnthropic
- EvaluatedAmazon Nova PremierAmazon
- IncumbentAmazon RekognitionAmazon
One labelled dataset built from the platform's own content. Every candidate invoked through Amazon Bedrock, scored on the same images, under the same prompts, in a Phase 1 evaluation deliverable — with Amazon Rekognition held as the incumbent router fallback.
“Bedrock wins. Bedrock is really the path, really clearly.”
The same discipline, three more ways
Claude 3.5 Sonnet on the reasoning steps, Claude Haiku on the latency-sensitive ones, benchmarked against Amazon Nova before launch — then hardened with Bedrock Guardrails against prompt injection.
Read the case study →Nova 2 Lite primary, Nova Pro fallback, Anthropic Claude qualified as the alternate foundation model on tool-call accuracy. Shipped 28 days early with zero cross-tenant incidents.
Read the case study →12,000 real production prompts mined into an automated evaluation suite with LLM-judge scoring. A head-to-head run caught a blind model switch before it reached customers.
Read the case study →How the model actually gets chosen.
Land on Bedrock
One governed model channel inside the customer's own AWS account. Every call metered, tagged and budget-enforced; guardrails and observability wired before the first prompt ships.
Build the labelled set
From the customer's own data, not a public leaderboard. On the moderation engagement that meant a labelled image set; on the PM SaaS it meant 12,000 real production prompts.
Score every candidate identically
Same data, same prompts, same harness — frontier labs, open weights and first-party models side by side. The result is a deliverable the customer keeps, not a slide.
Ship the winner, qualify the alternate
The winner takes the workload; a second model is qualified as the tested fallback. Because everything runs through one channel, changing your mind later is a config change with a receipt.
The estate, one column per provider.
Claude Partner Network member. Where reasoning quality and guardrail maturity carry the workload.
OpenAI Partner Network member since July 2026 — referral, co-sell and delivery services.
Qwen, Llama, Mistral and friends — in production wherever they win the benchmark outright.
Amazon Nova and Titan carry the default tier in two published production builds and appear in the head-to-head runs; Google Gemini is metered through the same gateway and evaluated where multimodal reach or context length decides it. They do not have their own pages here — they show up inside the engagements above, which is where the evidence is.