Back to all testimonials
Industry
Education & HR Tech
Media & Advertising
Services

Model Evaluation

Read the case study

Nonsense: LLM evaluation harness on AWS Bedrock

"Nonsense teaches languages through real Hollywood movie dialogue — per-phrase translation and contextual explanations, all generated by an LLM pipeline. Our models run on Gemini while our infrastructure runs on AWS, and we wanted an independent look at whether moving those workloads to Bedrock was worth doing.

Tech 42 built exactly the harness we needed, and built it well: Bedrock batch inference across three candidate models and nine languages, a full validation stack over our production JSON schemas, LLM-as-judge scoring, and a dashboard that made per-model and per-language results easy to dig into.

The team was a pleasure to work with. They moved fast when timing mattered, and stayed genuinely responsive to every concern we raised along the way. The final presentation was well-structured and easy-to-digest, besides also being a very nifty tech demo. 5 stars, would recommend!"

Nathan Glenn

NLP/Backend Software Engineer

Project summary

Nonsense teaches languages through Hollywood movie dialogue. The app translates phrases in real time and gives contextual explanations as users watch. Nonsense ran its core AI workloads on Gemini while its infrastructure sat on AWS. The team wanted an independent evaluation of moving those workloads to Bedrock. Tech 42 built a validation dataset from Nonsense's production data. The evaluation used Bedrock batch inference across three candidate models: Nova Pro, Claude Haiku 4.5, and Nova Lite 2. Testing covered nine languages and Nonsense's production JSON schemas, scored with an LLM-as-judge framework. Tech 42 documented performance metrics, including time to first token, latency, and tokens per second, and completed a cost analysis based on Bedrock Batch pricing. Results were delivered in a Jupyter notebook with README documentation, a sample Bedrock Batch job submission notebook, and a dashboard for comparing results by model and language.

Back to all testimonials