sinter.bio helps organisations scope, evaluate, validate and monitor clinical AI, healthcare LLMs, OMOP/FHIR and other health-data systems. We turn evaluation into a practical evidence pathway: intended use, benchmarks, safety, equity, implementation and continuous learning.
We support independent evaluations, formative assurance and post-deployment evidence generation for high-stakes digital health systems.
Each engagement is assembled around the claim being made, the risk profile, the intended users and the evidence needed for a credible decision.
Clarify intended use, risk tier, claims, evidence gaps, data access, governance and the right evaluation pathway before large costs are incurred.
Benchmark design, expert adjudication, safety and equity evaluation, RAG/LLM evals, interoperability and OMOP/FHIR assurance.
Version-aware monitoring, release checks, failure capture, decision trace and iterative evidence generation after deployment.
A clear pathway helps clients understand what happens next, what they receive at each step, and how trust and evidence deepen over time.
Complete the structured onboarding form so we can understand the product, setting, claims, evidence assets and governance constraints.
We assess evaluability, risk, likely modules, missing evidence and fit. Where appropriate, we compare the opportunity to anonymised patterns from prior engagements.
We refine the questions, define deliverables, propose a budget/timeline and assemble the smallest credible multidisciplinary team for the assignment.
Depending on the assignment, this may include benchmark design, expert workshops, human–AI studies, safety or equity analysis, implementation work, dashboards and release assurance.
We document provenance, decisions and lessons so findings remain reusable for future releases, procurement, publication and real-world improvement.
For suitable enquiries, the intake is more than an administrative form. It becomes the basis of an early strategic readout: where the organisation stands, which evidence is missing, which evaluation pathway is proportionate, and how similar cases have approached related challenges.
We do not promise spurious “industry rankings” off a tiny sample. Early comparative insights are framed carefully: trend observations, similar-case lessons and maturity signals, always anonymised and only when the internal reference set is credible enough to support them.
Evidence should help organisations decide what to do next — not simply produce a score.
An evaluation produces more than a score. We capture the context needed to interpret the evidence later: what was tested, which version, in which setting, against what reference standard, by whom, with which thresholds, what failed, what changed and why.
Over time, these structured traces can form an assurance context layer linking evidence, decisions, versions, settings and outcomes. That gives client teams an auditable organisational memory for re-evaluation, release decisions, incident review and handover.
Evidence becomes more valuable when its provenance and decision context remain visible.
Link each claim and decision to the cases, methods and evidence used to test it.
Record model, prompt, vocabulary, workflow or pipeline version alongside setting and user group.
Preserve reference standards, adjudication paths, metrics, thresholds and rationale.
Confirmed failures can become governed regression cases for future releases.
Repeated evaluation metadata can become a controlled vocabulary: stable terms for context, provenance, mapping quality, evidence status and decisions. Over time, that vocabulary can support an ontology and a context graph.
Where local terms, codes, languages, cadres or workflows are not represented cleanly, preserve them as governed source terms, map them to standard concepts where semantic equivalence exists, and retain the provenance, ambiguity and adjudication needed to interpret that mapping later.
R&D direction: modular African source-term and context vocabularies organised by country, language and domain, with reusable mapping provenance.
Source term, language, country, setting, standard concept, mapping confidence, reviewer, rationale, version and adjudication status.
Stable identifiers, definitions, synonyms, allowed values and change history.
Describe how terms, cases, evidence, failures and decisions relate.
Link datasets, benchmarks, settings, reviewers, versions, decisions, failures and outcomes.
We want African evidence to shape global health R&D early enough to influence what gets built, how it is evaluated and what claims are supportable.
sinter.bio has a particular interest in responsibly bringing African clinical data, local contextual knowledge, community priorities and implementation evidence into global health R&D and AI evaluation pipelines.
Evaluation should remain sensitive to language, workflow, infrastructure, disease patterns, population representation and health-system realities. Where community-governed or culturally specific knowledge is relevant, authority, provenance, responsibility, ethics and benefit are explicit parts of the design.
Clinical data-assurance work informing reusable methods for semantic mapping quality, terminology provenance and regression-oriented evaluation.
Current work with OMOP Bridge is helping shape a reproducible evaluation approach for AI-assisted clinical terminology mapping and data transformation. The methodology combines the OHDSI Usagi tool as a comparator with independent expert reference, clinical adjudication, benchmark design, failure taxonomy and regression-oriented testing.
A natural next layer is a governed source-term and mapping-provenance registry: preserving local terminology, country and language context, mapping rationale, ambiguity, reviewer decisions and vocabulary versions while linking to standard OMOP concepts where appropriate.
This case provides an early laboratory for the broader sinter.bio method system: benchmark discipline, evidence provenance, context-aware vocabularies, decision trace and lifecycle assurance.
After triage, sinter.bio scopes the disciplines needed for the claim and risk profile, then assembles the smallest credible project team from relevant specialists, subject to availability, independence and conflict checks.
In metallurgy, sintering is the process of forming a stronger material by applying heat or pressure so that particles bond and structure emerges — without simply melting everything into one indistinguishable mass.
That idea captures what sinter.bio is trying to do in clinical AI and health-data assurance. We bring together evidence, context, provenance and local knowledge to produce something stronger and more useful for decision-making, while preserving the identity and meaning of the underlying parts.
Preserve local meaning. Strengthen global usability. Keep the trace of how evidence became a decision.
Physician, implementation researcher, research ethicist and clinical-data/AI assurance practitioner. Kevin’s work spans clinical AI evaluation, AI governance, implementation science, OMOP/OHDSI, health-data interoperability and context-aware evaluation in African health systems.
Start an evaluation enquiry or use the downloadable overview deck for a short introduction.
No. Scope is based on intended use, claims, risk, available evidence and the decision that the evaluation needs to support.
Yes. Engagements can cover evaluability, benchmark and validation work before deployment, as well as production monitoring, incident learning and release-regression assurance after deployment.
Client data and confidential cases are not pooled or reused by default. Any cross-engagement reuse requires explicit rights and appropriate governance.
Yes. The assurance framework includes terminology mapping, ETL/data transformation, interoperability, provenance, data-quality and related health-data evaluation.
Complete the intake and we will determine the appropriate next step: strategic readout, evaluability review, protocol development, benchmark design, independent evaluation or ongoing assurance.