Healthcare Data Types

Published

Jun 2026

  • ID: CMDS-002
  • Type: Core
  • Audience: Clinical data learners, analysts, health researchers, and CDI system builders
  • Theme: Clinical data systems begin by understanding the kinds of healthcare data that describe patients, care, measurements, and outcomes

Clinical data does not come as one clean table.

It comes from many parts of the healthcare system.

A patient may have demographic records, clinic visits, diagnoses, procedures, medications, laboratory tests, vital signs, imaging reports, discharge summaries, referrals, insurance claims, and follow-up outcomes.

Each data type answers a different question.

Each data type has different strengths.

Each data type has different risks.

Before cleaning, modeling, or reporting clinical data, we need to understand what kind of healthcare data we are working with.

This chapter introduces the major healthcare data types used in clinical and medical data systems.


Why Healthcare Data Types Matter

In many analytical projects, people begin by asking:

What variables do we have?

In clinical data systems, a better first question is:

What kind of healthcare process created this data?

This matters because clinical data is not created mainly for research.

It is often created during care delivery, billing, monitoring, documentation, or reporting.

A diagnosis code may reflect a clinician’s assessment.

A medication order may reflect an intended treatment.

A laboratory result may reflect a biological measurement at a specific time.

An encounter record may reflect contact with the healthcare system.

An outcome field may reflect what was observed, documented, or followed up.

These are not interchangeable.

A clinical data system must preserve these distinctions.


A Simple Map of Healthcare Data

Healthcare data can be organized into several broad groups:

  1. patient identity and demographics,
  2. encounters and care episodes,
  3. diagnoses and clinical conditions,
  4. procedures and interventions,
  5. medications and treatments,
  6. laboratory and vital sign measurements,
  7. clinical notes and narrative records,
  8. imaging and diagnostic reports,
  9. outcomes and follow-up data,
  10. administrative, billing, and claims data.

These categories are not perfect.

Real systems often overlap.

For example, a hospital admission may include diagnoses, procedures, medications, labs, notes, imaging, and discharge outcomes.

But this map gives us a practical starting point.


Patient Identity and Demographics

Patient demographic data describes who the patient is.

Common fields include:

  • patient identifier,
  • age or date of birth,
  • sex or gender,
  • residence or region,
  • marital status,
  • occupation,
  • insurance category,
  • facility or clinic identifier.

Demographics are often used to describe a study population.

They are also used for stratification, adjustment, equity analysis, and risk modeling.

For example, age may be used to compare outcomes between younger and older patients.

Sex may be used to examine whether disease patterns differ across groups.

Region may help identify geographic differences in access or outcomes.

However, demographic variables require care.

They may be incomplete.

They may be inconsistently recorded.

They may be sensitive.

They may also reflect social, administrative, or structural categories rather than purely biological ones.

A clinical data system should treat demographic variables as important but not neutral.


Encounters and Care Episodes

Encounter data describes contact between a patient and the healthcare system.

An encounter may be:

  • an outpatient clinic visit,
  • an emergency department visit,
  • an inpatient admission,
  • a laboratory visit,
  • a pharmacy visit,
  • a telemedicine consultation,
  • a follow-up appointment.

Encounter records are important because they define time and context.

They help answer questions such as:

  • When did the patient receive care?
  • Where was the patient seen?
  • Was the patient outpatient or inpatient?
  • How many visits did the patient have?
  • Was the patient followed after treatment?

Encounter data often provides the backbone of clinical analysis.

Many clinical cohorts are built around index encounters.

For example, a study may define the first hospital admission for heart failure as the index encounter.

From there, analysts may examine prior diagnoses, medications during admission, lab results, discharge status, and readmission.


Diagnoses and Clinical Conditions

Diagnosis data records clinical conditions assigned to a patient.

These may come from:

  • clinician documentation,
  • coded diagnosis fields,
  • discharge summaries,
  • problem lists,
  • registry forms,
  • claims records.

Diagnosis data can support disease surveillance, cohort identification, risk adjustment, and outcome analysis.

Examples include:

  • diabetes,
  • hypertension,
  • tuberculosis,
  • HIV,
  • chronic kidney disease,
  • pneumonia,
  • sepsis,
  • pregnancy-related conditions.

Diagnosis data is powerful, but it can be imperfect.

A diagnosis code may not always mean confirmed disease.

It may indicate suspected disease, rule-out diagnosis, billing classification, historical condition, or active clinical problem.

Clinical data systems should therefore ask:

What does this diagnosis field actually represent in this system?

This question is especially important when using diagnosis data to define cohorts.


Procedures and Interventions

Procedure data describes actions performed for or on a patient.

Examples include:

  • surgery,
  • cesarean section,
  • dialysis,
  • imaging procedures,
  • laboratory procedures,
  • vaccination,
  • catheter placement,
  • wound care,
  • screening procedures.

Procedure data helps describe what care was delivered.

It can also define exposures.

For example, a study may compare patients who received a procedure with those who did not.

Procedure data may also be used to measure service utilization.

However, procedure records may depend heavily on local documentation and coding practices.

Some systems capture procedures in structured fields.

Others capture them only in clinical notes.

Some record ordered procedures, while others record completed procedures.

That difference matters.


Medications and Treatments

Medication data describes drugs prescribed, ordered, dispensed, administered, or reported by the patient.

These are different events.

A prescription does not always mean the patient took the medicine.

A medication order does not always mean the drug was administered.

A pharmacy dispensing record does not always prove adherence.

Medication data may include:

  • drug name,
  • dose,
  • route,
  • frequency,
  • start date,
  • stop date,
  • prescribing clinician,
  • dispensing date,
  • administration time.

Medication data is central to treatment analysis.

It can help answer:

  • What treatment did the patient receive?
  • When did treatment start?
  • Was treatment changed?
  • Was the patient exposed before an outcome occurred?
  • Were there possible drug safety concerns?

Clinical systems should distinguish medication intention from medication exposure.

That distinction protects interpretation.


Laboratory Results

Laboratory data contains measured biological results.

Examples include:

  • hemoglobin,
  • creatinine,
  • blood glucose,
  • viral load,
  • CD4 count,
  • liver enzymes,
  • malaria test results,
  • culture results,
  • pregnancy tests.

Laboratory data usually has several key components:

  • test name,
  • result value,
  • unit,
  • reference range,
  • specimen date,
  • result date,
  • abnormal flag,
  • facility or laboratory source.

Laboratory data is highly valuable because it can provide objective clinical measurements.

But laboratory data also requires careful cleaning.

The same test may appear under different names.

Units may differ.

Values may include text such as positive, negative, <5, or not detected.

Dates may reflect collection time, processing time, or reporting time.

A clinical data system must standardize laboratory results before analysis.


Vital Signs and Bedside Measurements

Vital signs are common clinical measurements collected during care.

Examples include:

  • temperature,
  • blood pressure,
  • heart rate,
  • respiratory rate,
  • oxygen saturation,
  • weight,
  • height,
  • body mass index.

Vital signs are often repeated over time.

This makes them useful for monitoring disease severity, treatment response, and clinical deterioration.

For example, oxygen saturation may help identify respiratory compromise.

Blood pressure may help monitor hypertension.

Weight may help track nutritional status or fluid changes.

Vital signs can also be messy.

They may contain impossible values.

They may be recorded with inconsistent units.

They may be copied forward.

They may be missing when patients are stable or when systems are overloaded.

Clinical data readiness requires range checks, unit checks, and time checks.


Clinical Notes and Narrative Data

Clinical notes are free-text records written by healthcare workers.

They may include:

  • history and physical examination notes,
  • progress notes,
  • nursing notes,
  • operative notes,
  • discharge summaries,
  • referral notes,
  • radiology narratives,
  • pathology narratives.

Clinical notes can contain rich information that structured fields miss.

They may describe symptoms, clinician reasoning, social context, disease severity, treatment plans, and patient preferences.

But notes are difficult to analyze.

They are unstructured.

They may contain abbreviations.

They may mix current and past information.

They may include sensitive details.

They may require natural language processing or careful manual abstraction.

For many clinical systems, notes are a valuable source of context, but they should not be treated as simple structured variables without validation.


Imaging and Diagnostic Reports

Imaging data includes both images and reports.

Examples include:

  • X-ray,
  • ultrasound,
  • CT scan,
  • MRI,
  • echocardiography,
  • pathology slides,
  • radiology reports.

For many clinical analyses, the report is more accessible than the image itself.

A radiology report may contain findings, impressions, recommendations, and diagnostic conclusions.

Imaging reports may help identify disease severity, complications, or outcomes.

However, imaging data also has challenges.

Reports may use non-standard language.

Findings may be uncertain.

Images may require specialized storage, labeling, and interpretation.

When imaging is used for machine learning, careful attention is needed to avoid leakage, biased labels, and poor generalization.


Outcomes and Follow-Up Data

Outcome data describes what happened to the patient.

Examples include:

  • recovery,
  • death,
  • discharge status,
  • readmission,
  • disease recurrence,
  • treatment failure,
  • adverse event,
  • loss to follow-up,
  • pregnancy outcome,
  • laboratory improvement,
  • symptom resolution.

Outcome data is central to clinical interpretation.

But outcomes are not always directly observed.

Some outcomes are recorded only if the patient returns.

Some are captured only during hospitalization.

Some are inferred from records.

Some are missing because follow-up was incomplete.

A clinical data system should always ask:

Was this outcome actively measured, passively recorded, or inferred from available data?

That question affects how confidently we can interpret results.


Administrative, Billing, and Claims Data

Administrative and claims data are created for operations, reimbursement, reporting, or insurance processes.

They may include:

  • facility visits,
  • billing codes,
  • insurance claims,
  • service dates,
  • diagnosis codes,
  • procedure codes,
  • payment information,
  • provider identifiers.

These data can cover large populations and long time periods.

They are useful for health services research, utilization analysis, cost analysis, and population-level monitoring.

However, administrative data is shaped by payment and reporting rules.

A code may reflect reimbursement needs as much as clinical reality.

Therefore, claims-based definitions often need validation.


Structured vs Unstructured Data

Healthcare data may be structured, semi-structured, or unstructured.

Structured data has defined fields.

Examples include diagnosis codes, lab values, medication orders, and encounter dates.

Semi-structured data has some organization but may still require parsing.

Examples include forms, templated notes, and mixed text-result fields.

Unstructured data is mostly free text or raw media.

Examples include clinical notes, scanned documents, images, and audio recordings.

The more unstructured the data, the more interpretation work is required.

But structured data is not automatically correct.

A clean-looking code can still be clinically misleading.


Time in Healthcare Data

Time is one of the most important dimensions in clinical data.

Clinical events happen in sequence.

A diagnosis may occur before treatment.

A lab test may occur before or after medication.

An outcome may occur after discharge.

A model may accidentally use future information if timing is not handled carefully.

Clinical data systems should track:

  • date of birth,
  • encounter dates,
  • diagnosis dates,
  • procedure dates,
  • medication start and stop dates,
  • specimen collection dates,
  • result dates,
  • outcome dates,
  • follow-up windows.

Temporal structure protects against false conclusions.

It also supports reproducible cohort definitions.


Data Type, Analytical Role, and Interpretation

The same variable can play different analytical roles.

For example, a laboratory result may be:

  • an inclusion criterion,
  • a baseline characteristic,
  • a predictor,
  • a monitoring variable,
  • an outcome,
  • a safety signal.

A diagnosis may be:

  • a cohort-defining condition,
  • a comorbidity,
  • an exclusion criterion,
  • an outcome,
  • a risk adjustment variable.

A medication may be:

  • an exposure,
  • a treatment group definition,
  • a confounder,
  • a marker of disease severity.

This means the analyst must define not only what a data field is, but how it is being used.

Clinical interpretation depends on analytical role.


Common Risks Across Healthcare Data Types

Several risks appear across many clinical data types:

  • missing data,
  • duplicate records,
  • inconsistent identifiers,
  • inconsistent units,
  • unclear dates,
  • coding variation,
  • copy-forward documentation,
  • incomplete follow-up,
  • facility-specific workflows,
  • sensitive personal information,
  • changes in clinical practice over time.

These risks do not mean the data is unusable.

They mean the data must be handled as clinical evidence, not just as spreadsheet content.


CDI Principle: Respect the Source Process

A core CDI principle for clinical data systems is:

Do not interpret a clinical variable without understanding the healthcare process that created it.

This principle applies to every data type.

A lab result comes from a test order, specimen collection, analysis, and reporting process.

A diagnosis comes from clinical assessment, documentation, and coding.

A medication record comes from ordering, dispensing, administration, or reporting.

An outcome comes from observation, follow-up, or inference.

When we respect the source process, our analysis becomes more defensible.


Practical Checklist

Before analyzing a clinical dataset, ask:

  • What healthcare data types are included?
  • Which tables describe patients?
  • Which tables describe encounters?
  • Which tables describe diagnoses, procedures, and medications?
  • Which tables contain measurements such as labs and vital signs?
  • Which fields define outcomes?
  • Are dates available for key clinical events?
  • Are values coded, numeric, text, or mixed?
  • Which variables are sensitive?
  • Which variables require clinical interpretation?
  • Which variables require standardization before analysis?

This checklist helps move from raw clinical records to analysis-ready thinking.


Chapter Summary

Healthcare data systems contain many different types of data.

Patient demographics describe who received care.

Encounters describe contact with the healthcare system.

Diagnoses describe clinical conditions.

Procedures and medications describe care delivery.

Laboratory results and vital signs describe measurements.

Clinical notes and imaging reports provide narrative and diagnostic context.

Outcomes describe what happened after care.

Administrative and claims data describe operational and billing processes.

Each data type has value.

Each data type has limitations.

A strong clinical data system begins by recognizing these differences before moving into cohort design, governance, cleaning, modeling, and reporting.