Back to all testimonials
Industry
Health & Life Sciences
Services

Model Evaluation

Read the case study

InferaHealth: Repeatable model evaluation framework on Amazon Bedrock

"Tech 42 helped us build a repeatable model evaluation framework for our AI workflows on AWS Bedrock. Using a validation dataset we provided from our own prompt/response data, Tech 42 benchmarked our current model against Claude Sonnet 4.6, Claude Haiku 4.6, and Gemma 3 using custom evaluation metrics covering our internal workflow's testing requirements along with latency and cost tracking per run. Every workflow in scope was covered and tested. We now have Bedrock fully integrated into our evaluation process, and the framework itself is far more consolidated than what we had before.

One of the biggest wins was having a fully customizable and scalable evaluation framework embedded in our own codebase, giving us far more flexibility than relying on Langfuse's built-in LLM-as-a-Judge evaluations. Consolidating our evaluation process and wiring it directly into Bedrock has meaningfully simplified how we test and compare models going forward.

Tech 42 delivered a working notebook for running model evaluations, along with a knowledge transfer session covering the results and recommendations, plus a recording for internal reference. We are already planning a second phase to extend evaluation coverage to our agentic chatbot and other workflows."

Theo Chitayat

Co-founder & COO

Project summary

InferaHealth is a healthcare AI company providing clinical decision support tools to hospitals and physician groups. Tech 42 built a repeatable model evaluation framework on Amazon Bedrock. The project began with a clinical validation dataset built from de-identified prompt and response data and clinician annotations. Tech 42 benchmarked InferaHealth's existing model against Claude Sonnet 4.6, Claude Haiku 4.6, and Gemma 3 on Amazon Bedrock. Evaluation covered conversation completeness, turn relevance, knowledge retention, latency, and cost per conversation, using the Deepeval framework and Confident AI for test observability. Every workflow in scope was evaluated and tested. Deliverables included a validation dataset, a Jupyter notebook for running evaluations, a comprehensive report with model recommendations, and README documentation. The work gave InferaHealth a consolidated, repeatable way to compare models on AWS Bedrock and evaluate migration options while maintaining HIPAA and BAA compliance.

Back to all testimonials