Verticalized Data & Environments Lab · Medical AI

We teach AI to reason like a physician

The most valuable medical knowledge isn't written down — it lives in how physicians reason. SinoHealth AI is a verticalized data & environments research lab for medical AI: we turn clinical judgment into evals, RLHF, de-identified datasets and reasoning environments — physician-in-the-loop and compliant. Singapore-based, serving medical AI worldwide.

Industries we power:
Healthcare & Medicine · NowLife Sciences · ExpandingFinance · ExpandingLaw · ExpandingScience & Code · Expanding
The Problem

Medical AI can answer. It can't yet reason like a clinician.

Real clinical judgment — the differentials, the follow-up questions, the safety instincts — isn't written down. It lives in how physicians think. We capture it, and turn it into data models can learn from.

Models trained on outputs plateau. Models trained on reasoning improve.

Our approach

Serious medical AI needs clinical judgment, not crowdsourcing

Leading LLMs do fine on clean case notes, but break on patients' real, incomplete accounts. Fixing that starts with data only practicing physicians can produce — measured, structured, and compliant.

Physicians in the loop

Every task is handled by attending-level specialists with verified credentials, drawn from a global clinical network — not generic labelers.

Reproducible quality

Dual annotation → senior adjudication → gold-set audit, with Cohen's kappa monitored throughout. Quality is measurable and traceable.

Compliant across jurisdictions

De-identification, ethics review and configurable data residency, aligned with HIPAA, GDPR and PDPA — compliance is built into the pipeline.

κ ≥ 0.6
Target annotation agreement
9+
Clinical specialties, expandable
8+
Languages delivered, and growing
3
HIPAA · GDPR · PDPA compliant
How it works

A reproducible pipeline, from raw data to model-ready

01 / Ingest

Source & de-identify

Ingest raw clinical data through compliant channels; de-identify and normalize to a clean, structured base.

02 / Enrich

Structure & annotate

Physicians structure, label, and build eval / RLHF data; dual annotation with senior adjudication and kappa.

03 / Deliver

QA & deliver

Gold-set audit and double-blind review, then delivery as JSONL + spec + QA report, in your language.

Our Flywheel

Research-first: a loop that compounds

We go beyond labeling: we identify where medical models fail, build the data and rubrics to fix it, verify the gains, and publish — a loop where every turn compounds the next.

1

Benchmark

Construct realistic, challenging clinical tasks that surface where models fail.

2

Data, rubrics & environments

Physicians build targeted data, rubrics and environments for each failure mode.

3

Verify

Post-train and verify the signal measurably improves models.

4

Leaderboard

Publish open results — credibility and inbound demand.

5

License data & evals

Model builders license the data and evaluations behind the results.

OPEN BENCHMARK · COMING SOON

ClinReason-Bench

The starting point of our flywheel — an open benchmark for physician-grade clinical reasoning and safety, spanning specialties and languages.

See a clinician-grade eval sample

Spec, scoring rubric and worked samples are ready to demo against your model's evaluation needs.

Solutions

Three layers: from data resource to intelligent application

We move medical data from raw resource, to standardized data product, to deployed application — physician-in-the-loop and compliant at every layer.

01 · Data

Data resources

Compliant access + de-identification + structuring of clinical, imaging and lab data (China & global sources).

Sourcing · de-identification · structuring
02 · Products

Data products

Standardized, documented datasets, eval sets and RLHF data — with transparent rights & provenance.

03 · Applications

Intelligent applications

Deploy clinical AI — e.g. a hospital consultation agent (software + hardware) — into partner hospitals.

Capabilities

The medical-data pipeline for AI models

End-to-end data pre-processing for medical AI — from raw records to model-ready datasets, every step delivered by practicing physicians.

Pre-processing

Data pre-processing & curation

De-identification, cleaning, structuring and normalization (ICD coding) of raw clinical data into model-ready datasets.

  • De-identification
  • Structuring & normalization
  • ICD-10 / ICD-9-CM-3
  • Quality control
Annotation

Expert annotation & labeling

Clinician-in-the-loop labeling across clinical text and imaging — entity, relation, diagnosis and reasoning labels.

  • Clinical NER & relations
  • Diagnosis labeling
  • Reasoning chains
  • Image annotation
RLHF

RLHF & preference data

Physicians score, rank and rewrite model outputs into clinically-grounded preference and SFT data.

  • Preference ranking
  • Rewrites & corrections
  • SFT demonstrations
  • Multi-turn consults
Evals

Evaluations & benchmarks

Diagnostic-reasoning, safety and guideline-alignment evals that quantify how a model actually performs.

  • Diagnostic reasoning
  • Guideline alignment
  • Six-dimension rubric
  • Custom benchmarks
Safety

Red-teaming & safety

Adversarial testing for medical safety — surfacing unsafe advice, hallucination and missed red flags before your users do.

  • Adversarial prompting
  • Red-flag detection
  • Hallucination checks
  • Over-refusal balance
Multimodal

Multimodal & imaging

Radiology and pathology specialists for image labeling, report-quality assessment and image-text alignment.

  • Lesion annotation
  • Report scoring
  • Image-text alignment
  • Structured extraction

Need a custom data pipeline?

Tell us your model, use case and evaluation goals; we'll return an actionable plan and quote.

Data Products

A catalog of model-ready medical data products

Not one-off labeling — standardized, documented data products with transparent rights and provenance. Every product states its source, processor and licensed use.

SHAI-EVAL-01

Clinical Diagnostic-Reasoning Eval Set

Use: model evaluation & benchmarking (information-seeking, differentials, safety)
6dimensions
κ≥.6agreement
9+specialties
ContentConsumer + specialty cases, standardized & incomplete-report modes, physician gold labels + rubric
EvalSafetyMultilingual
Physician-built · de-identifiedAvailable
SHAI-EMR-02

Full-Lifecycle EMR Dataset (de-identified)

Use: pre-training corpus / fine-tuning / RAG knowledge base
10+record types
ICDcoded
OP + IPoutpatient + inpatient
ContentAdmission, progress notes, orders, labs, imaging reports, discharge — with original full text, longitudinally linked
Pre-trainingFine-tuneRAG
Licensed channel · de-identifiedBuildable
SHAI-RLHF-03

Physician Preference & RLHF Dataset

Use: alignment — preference optimization & SFT
2×dual-rated
SFT+ demos
Rxsafety
ContentPhysician scoring, ranking and rewrites of model outputs; multi-turn consults; safety red-lines
RLHFPreferenceSFT
Physician-built · de-identifiedBuildable
SHAI-IMG-04

Medical Imaging + Report Dataset

Use: multimodal training & report generation eval
WSIpathology
DICOMradiology
+Rptreports
ContentPathology WSI & radiology studies with findings, conclusions and diagnosis labels; image–text aligned
MultimodalLesion labels
Global sources · consent-tracedBuildable
SHAI-TRACE-05

Retrieval & Reasoning-Trace Dataset (traceable)

Use: RAG-agent training & pre-launch step verification
↩source-linked
Stepby step
QAreviewed
ContentQuery → evidence selection → summary; every claim's citation maps back to the source text span (character offsets)
RAGTraceableAgent
Licensed channel · de-identifiedBuildable
SHAI-ENV-06

Clinical Reasoning Environment (RL)

Use: RL post-training & agent evaluation in a simulated consultation
↺resettable
Autoverifier
SP+ real reports
ContentA resettable, auto-scored virtual consultation — standardized patients + real incomplete self-reports; physician-built safety & reasoning verifiers
RLEnvironmentAgent
Physician-built verifiersRoadmap
Custom

Need a custom data product?

Tell us your model, task and target metrics — we'll scope a product with a full spec, rights statement and quote.

Every product ships with a data dictionary and a rights & provenance statement (source institution · processor · licensed use). Figures are illustrative and confirmed at scoping.

Who we serve

Built for every buyer of medical data

Different teams need different things from clinical data. We tailor products and services to each.

Medical AI

Medical-AI & LLM companies

Evals, RLHF, de-identified corpora and fine-tuning data to make models reason and stay safe.

Needs: reasoning evals · RLHF · pre-training corpus
Pharma & Device

Pharma & medical-device

De-identified real-world & imaging datasets for research, biomarker and model development.

Needs: RWD · imaging · cohorts
Hospitals

Hospitals & medical schools

Co-develop a dedicated consultation agent (software + hardware) that lifts physician efficiency.

Needs: clinical agent · research collaboration
Insurers

Insurers & health platforms

Structured clinical data and evals for underwriting, claims and health-management models.

Needs: structured data · risk models
Specialties

Specialties we cover, expandable on demand

Each specialty is staffed with practicing specialists from our global network and matched to high-value tasks.

On

Oncology

Diagnostic-reasoning & treatment eval, guideline alignment, staging and drug checks.

Rx

Imaging / Radiology

Lesion annotation, report-quality scoring, image-text alignment.

Cv

Cardiology

Red-flag detection, differentials, ECG-imaging combined eval.

En

Endocrine / Chronic

Chronic-care pathways, family-doctor eval, medication labeling.

Pu

Respiratory

Information-seeking checks, differential correction, imaging-symptom integration.

Pd

Pediatrics

Pediatric disease labeling, dosage safety, growth assessment.

Nr

Neurology

Neuro differentials, imaging annotation, scale-based assessment.

Pa

Pathology

Pathology image labeling, diagnostic-concordance scoring, structured reports.

Em

General / Emergency

End-to-end clinical decision support, red-flag safety eval, triage pathways.

Quality & Compliance

Measurable quality, compliant by design

From physician credentials to delivery, every step has standards and metrics — and data is compliant across the jurisdictions you operate in.

01

Credential verification

Attending-level and above; licenses and specialty backgrounds are strictly verified before onboarding.

02

Dual independent annotation

Each record is annotated independently by two physicians to avoid single-rater bias.

03

Senior adjudication

Disagreements escalate to senior experts, forming the gold version.

04

Agreement monitoring

Cohen's / Fleiss' kappa is monitored throughout; below threshold triggers a calibration session.

05

Gold-set audit

Known-answer items seeded per batch; ≥20% double-blind review before delivery.

06

Data governance

De-identification, ethics/IRB review and configurable data residency across jurisdictions.

Compliance & security

HIPAAUS health data
GDPREU data protection
PDPASingapore
ISO 27001Infosec (in progress)
SOC 2On roadmap
IRBEthics review

De-identification, DPAs, and data-residency options are configured per client and jurisdiction.

For Physicians

Put your clinical expertise into the next generation of medical AI

Join SinoHealth AI's global network of licensed physician experts — remote, flexible, paid per project — and shape the evaluation and training data behind frontier medical AI.

Benefits

Why join the SinoHealth AI expert network

Remote & flexible

Work online on your own schedule, alongside clinical practice.

Paid per project

Paid by expert-hour; scarcer specialties earn more.

Frontier work

Directly shape evaluation and alignment for leading medical AI.

Use your specialty

Define the clinical bar for AI with your specialist judgment.

We're looking for

Physicians we're looking for

  • ✓A valid medical license
  • ✓Attending-level or above (some projects accept senior residents)
  • ✓A clear specialty
  • ✓Rigorous, evidence-based clinical thinking

How it works

01

Apply with credentials

Fill in the form with your specialty and practice details.

02

Verification & trial

We verify your license and you complete a short trial task.

03

Get matched & earn

Get matched to projects by specialty, paid by the hour.

Apply to the expert network

We'll contact you after verification. * required.

Your application is emailed to us; we'll reach out after verification.

For Hospitals · Research Collaboration

A dedicated consultation agent, built for your hospital

In collaboration with medical schools and hospitals in China and abroad, we combine their clinical resources with our data to research and build a hospital-specific consultation agent — software and hardware — that helps physicians work faster.

What we build together

Clinical resources × our data → a hospital's own agent

We co-develop with each partner: their specialists and research capacity, our de-identified data and physician-annotation methodology, delivered as a system the hospital owns.

Medical-school resources

Clinical expertise and research capacity from partner medical schools and their affiliated hospitals.

Our data foundation

De-identified, multi-source clinical data and reproducible physician annotation — the fuel a reliable agent needs.

Software + hardware

A consultation agent tuned to the hospital's specialties and workflow, delivered as an integrated software-and-hardware system.

How it helps physicians

Less time on the routine, more on the patient

Pre-consultation intake

Structured history-taking before the visit, so physicians start with a clear picture.

Information-seeking

The agent asks the right clarifying questions instead of jumping to conclusions.

Documentation

Drafts notes and summaries, cutting time spent on paperwork.

Decision support

Guideline-aligned suggestions with safety and red-flag checks, physician-in-command.

Hospital-specific

Tuned to the hospital's departments, pathways and language.

Physician-in-command

Assists, never replaces — the clinician stays in control at every step.

Exploring a hospital or medical-school collaboration?

We're in discussions with several institutions in China and abroad. Tell us about yours.

About

A verticalized data & environments research lab for medical AI

SinoHealth AI is a Singapore-based verticalized data & environments research lab for medical AI, serving model builders worldwide. We pair SinoHealth's clinical network with a global roster of expert physicians.

We believe the ceiling of medical AI is set by how much real clinical judgment lives in its data. So we embed physicians in every step — sourcing, structuring, annotation, RLHF, evaluation and reasoning environments — with reproducible QA and compliance built in.

·
Global physician network generic crowdsourcing can't reach
·
Combined clinical + ML-eval methodology; measurable data
·
Compliant across jurisdictions: HIPAA, GDPR, PDPA
·
Multilingual delivery for models going global
"The gap in medical AI is rarely parameters — it's whether a physician's judgment is in the data."
— SinoHealth AI
HQ
Singapore · serving worldwide
Focus
Data & environments for medical AI
Delivery
Data · Annotation · RLHF · Evals · Environments
Languages
ENFRDEPTRUESARZH
Get in touch

Book a demo, or request a sample

Leave your details and we'll share a clinician-grade eval sample and a one-pager, then set up a 20-minute call.

Submissions are emailed to bryson@sinohealth.ai; we'll reply shortly.

Email
bryson@sinohealth.ai
WhatsApp
+852 6218 3814
HQ
Singapore