/ Service

LLM evaluation and retrieval quality

Without an evaluation set, retrieval tuning is opinion and every change is a gamble. I build the measurement first — what counts as a correct, grounded answer — so changes to chunking, search and reranking become comparable engineering decisions.

/ What this covers

  • Evaluation sets and scoring for grounded, accurate answers
  • Retrieval tuning across chunking, hybrid search, reranking and vector retrieval
  • Evaluation workflows wired into delivery pipelines so regressions surface early
  • Qualitative architecture evaluation of AI frameworks and platforms, recorded as ADRs
  • Synthetic load testing of gateway and serving paths, reported as synthetic and kept separate from production performance claims

/ Evidence

Retrieval architecture · ReAct agents

Enterprise RAG & Agent Platform

Architected and delivered a scalable enterprise RAG and agent platform: a 500GB+ embedding and ingestion pipeline feeding hybrid search with reranking and vector retrieval, behind a ReAct agent architecture. Retrieval quality was treated as the product — measured chatbot accuracy improved to 95% and response latency was reduced.

Read the Enterprise RAG & Agent Platform case study

Grounded assistant · Procurement

Enterprise Procurement Knowledge Assistant

Designed and productionised a grounded enterprise Procurement assistant on Azure AI Foundry, using LangGraph for orchestration, ChatKit for the interface, PGVector for retrieval and integrations with MCP servers. Answers stay grounded in trusted enterprise documents, and the architecture was standardised into a reusable reference implementation.

Read the Enterprise Procurement Knowledge Assistant case study

/ Related services

  • Agentic AI engineering

    ReAct agent architectures and grounded assistants, with tool access, guardrails and evaluation designed in.

  • Enterprise AI platforms

    Paved paths and internal platforms so AI applications start from a supportable, governed baseline.

  • MCP & tool integrations

    Governed agent access to internal systems through MCP servers and shared integration patterns.

  • AI governance

    The operating model — identity, approvals, usage and cost controls — that lets AI scale past pilots.

/ Next step

Let's talk about the work.

Open to full-time Senior–Staff AI / agent platform roles (remote-friendly), as well as contract and consulting engagements. Response within two business days.