Computer Vision in Healthcare: Use Cases, Tools, and Real Examples for 2026
A phone camera can watch someone do a squat and tell them their knee is collapsing inward. A CT scan can be triaged for a stroke before a radiologist has opened it. Both run on computer vision, the same family of models pointed at very different problems.
Computer vision in healthcare raises four questions worth answering in one place: what it does across medical specialties, which tools are worth building on, what the FDA does with a feature that reads medical images, and what a build costs. Answers here stay at survey depth, and three sibling guides go deeper. Pose estimation is the technique deep-dive, skin cancer detection walks one application build end to end, and AI-assisted medical image documentation covers the reporting workflow downstream of the model.
What is computer vision in healthcare used for?
Reading medical images is the bulk of it: 1,230 of the 1,614 AI-enabled devices on the FDA’s list went through the radiology panel. Outside radiology, computer vision screens for diabetic retinopathy with no physician in the loop, flags suspicious regions on pathology slides, totals surgical blood loss from photographs of sponges, and measures joint angles from a patient’s phone video at home.
Key takeaways
- The FDA’s own list of AI-enabled medical devices held 1,614 authorizations when the agency last refreshed it in September 2026. Radiology accounts for 1,230 of them, and yearly authorizations went from 80 in 2019 to 335 in 2025.
- Vision-language models changed what a build starts from. MedSAM cut expert annotation time by about 82% in its own controlled study, while general-purpose models like GPT-5 and Gemini still miss findings at rates no clinical product could ship on.
- A feature that analyzes a medical image for a clinical purpose is a medical device. FDA’s clinical decision support guidance closes that exemption explicitly, and 96% of AI devices reach the US market through 510(k).
- Our own AI-assisted healthcare builds ran $25,000 to $97,000 with a $49,000 median, and a production custom ML integration runs $70,000 to $200,000 and up on top of that.
Table of Contents:
- What is computer vision in medicine
- Vision-language models changed what a medical CV build starts from
- What computer vision in the medical field actually does
- Main advantages of computer vision in healthcare
- Computer vision healthcare use cases by medical specialty
- A computer vision feature is what turns your app into a medical device
- Which tools you build on, and which ones can touch PHI
- What a medical computer vision build actually costs
- What we learned building a computer vision RTM platform
What is computer vision in medicine
Computer vision is the technology that lets a computer, whether that’s a smartphone, a pair of smart glasses or a medical device, recognize images and video and make sense of what they contain. Computer vision in the medical field turns a scan, a slide, a retina photo or a video of someone moving into something measurable.
Medical computer vision is machine learning algorithms trained to recognize particular objects and their attributes in a medical image. Convolutional neural networks carried that work for a decade: gather enough labelled examples of a finding, then train a model to spot it. One model, one job.

The output side is what decides whether you are building a medical device.
That one-model-one-job pattern is still how most shipped medical image analysis AI works, and it’s the part that has changed hardest since we first wrote this page. If you’re weighing a build, our guide to machine learning app development walks the whole path.
For one number on how real any of this is, take the regulator’s. The FDA publishes a list of the AI-enabled medical devices it has authorized for the US market, and when the agency last refreshed that page in September 2026 it held 1,614 entries, against 80 authorizations in the whole of 2019 and 335 in 2025 alone. Radiology accounts for 1,230 of them. FDA cautions that the list isn’t comprehensive, so read it as a floor.
Market forecasts for medical computer vision applications are worth a good deal less than that. Seven research firms published estimates inside the last 16 months, and they disagree about 2025 by more than threefold, from $1.55 billion to $4.86 billion, with growth rates running from 24% to 49%. One of them prints $2.01 trillion for AI in medical imaging on a page it updated the same day as a $3.51 billion computer vision page. This article carried a number of that kind for years. We took it out.
Related: medical device integration and medical device cost breakdown
Vision-language models changed what a medical CV build starts from
Everything above was built the old way, one labelled dataset and one trained model per task. Foundation models moved the starting line. They also picked up a set of claims that don’t survive contact with the published evidence, and on a clinical topic that gap is the whole story.
What general-purpose multimodal AI does with a scan
Ask GPT-5, Gemini or Claude to read a chest X-ray and a confident paragraph comes back. The peer-reviewed record on whether that paragraph is correct is consistent and unflattering.
A 2024 study in Radiology ran 515 images from 470 patients through GPT-4V. The model named the imaging modality correctly 100% of the time and the anatomic region 99.2% of the time. Then it reached the actual job: free-text diagnostic accuracy ran from 0% on pneumothorax to 90% on brain tumours, specificity fell to 8% for brain haemorrhage against human readers at 97.2%, and 86.5% of its free-text reports contained a false positive. The authors concluded that GPT-4V “failed to detect, classify, or rule out abnormalities in image interpretation.”
The 2026 generation hasn’t closed that gap. A February 2026 study in Diagnostics put 65 fracture-positive hand radiographs through 4 current models, 5 runs each: GPT-5 Pro reached 64.3% accuracy, Gemini 2.5 Pro 56.9%, Claude Sonnet 4.5 33.8%. The trap those authors flag is the one that catches product teams. The most reproducible model in the set was also among the least accurate, so agreement between runs tells you nothing about correctness.
Where multimodal AI earns its keep is the metadata around the image. Modality identification at 100% across 2 independent studies, anatomic region between 87% and 99%. Accuracy also climbs when text context travels with the image. That points a build at routing studies, protocolling, drafting the boilerplate half of a report and pulling structure out of prior text. Our guide to AI-assisted medical image documentation covers that workflow in build detail.

The gap between those two groups is the whole build decision.
MedGemma and MedSAM are the parts you can build on
Two model families are doing real work here, and neither one is an API you call.
MedGemma is Google’s open-weight medical vision-language model, built on Gemma 3. You download the weights and host them yourself, which keeps images inside your own perimeter. The current release, MedGemma 1.5, handles chest X-ray, CT and MRI volumes, whole-slide pathology, dermatology and fundus images. Read two things before scoping anything on it. The licence is Google’s own Health AI Developer Foundations terms, not Apache or MIT. And Google’s model card states the outputs “are not intended to directly inform clinical diagnosis, patient management decisions, treatment recommendations, or any other direct clinical practice applications.” In Google’s own expert review of generated chest X-ray reports, 49% of those covering abnormal studies were judged equal to or better than the radiologist’s original.
MedSAM is the segmentation story, and of the two families it has the harder evidence behind it. Published in Nature Communications in January 2024, it’s Meta’s Segment Anything model fine-tuned on 1,570,263 medical image-mask pairs across 10 modalities. Across 146 validation tasks it matched or beat specialist U-Nets trained per modality, and it held accuracy on external datasets where those specialist models fell away. Stock Segment Anything scored worst on most medical segmentation tasks in the same benchmark. Stock Segment Anything does well on endoscopy, where the pictures resemble the natural photographs it trained on.
What a foundation model removes is the per-task training run. The per-image human effort is still there: MedSAM is promptable and two-dimensional, so it wants a bounding box per target per slice, with 3D volumes handled as stacked slices. Its authors name 2 limits outright: a training mix skewed toward CT, MRI and endoscopy, and trouble with vessel-like branching structures where a bounding box is ambiguous.
What foundation models do to your budget
The pitch here runs well ahead of what anyone has measured. Two inputs have real numbers behind them. In MedSAM’s own controlled study, two abdominal radiologists annotated 733 tumour slices both ways, and model-assisted annotation cut their time by about 82%. Google reports its MedSigLIP encoder reaching strong chest X-ray classification from 512 or more labelled examples. On end-to-end build cost and timeline, nothing has been published at all.
One thing to hold onto before this turns into a roadmap. None of it sits in a cleared device. FDA’s device list, current through June 2026, still describes tagging foundation-model devices as something the agency “will explore,” and its August 2026 paper on generative AI devices says outright that it “does not represent draft or final guidance.” Our write-up of the first FDA-cleared LLM clinical AI walks the one 510(k) that made it through, and it’s a text tool rather than an imaging one.
What computer vision in the medical field actually does
Applications of computer vision in the medical field split two ways: clinicians using it to treat patients better, and patients using it on their own phones.

The clinician column is where the regulated products are.
What clinicians use computer vision for
These are the jobs it already does inside hospitals and clinics.
Diagnosing
Radiology AI does more of this work than every other specialty combined. Models read CT, MRI, ultrasound and X-ray studies at a scale no reading room can match, and they pick up patterns that are genuinely hard to see by eye, which moves diagnosis earlier. It’s a standing part of our healthcare app development services.
Example: Viz.ai’s triage software watches CT angiography for a large vessel occlusion and pages the on-call specialist directly. In the data behind its 2018 De Novo authorization, those automated alerts reached the specialist an average of 52 minutes sooner than the standard workflow, with individual cases ranging from 6 to 206 minutes. FDA created a device classification for triage software to let it through.
Reports automation
Reports automation falls out of the analysis computer vision algorithms already produce, and it’s a big part of why hospitals buy the analysis at all. Once the imagery has been read, computer vision-powered software pushes the measurements and the examination results into the systems that already hold the patient record, without anyone retyping them.
Patient condition assessment
Computer vision in medicine also gets pointed at the patient: how much blood they’ve lost, how alert they are, whether they’re in pain. You train the software on the physical and mental states you care about, and it watches for them while your clinicians work.
Gauss Surgical’s Triton runs on an iPad in the operating room, photographs surgical sponges, and totals blood loss while the procedure is still going. FDA cleared the sponge system in 2017, and Stryker now ships the capability as SurgiCount+.
Clinical trials
Computer vision-enabled mobile apps let sponsors and sites confirm on camera that a participant actually took the dose. That takes out the pill count at the next site visit and the coordinator hours attached to it.
Example: AiCure’s H.Code platform watches through the participant’s own phone camera to verify a dose was actually taken, and sells to sponsors, CROs, clinical sites and providers. Worth knowing where the line sits here: the consumer “pill identifier” tools from Drugs.com and WebMD match on an imprint code, colour and shape that you type in, then return a stock photograph, so the recognition work stays yours.
Read more on clinical trial software development
Surgery assistance
Put computer vision and AR into a surgeon’s smart glasses and the guidance lands in the field of view during the procedure. The heavier payoff is what happens to the recorded video afterwards.
Example: Touch Surgery, which Medtronic picked up when it acquired Digital Surgery in 2020, records surgical video, blurs identifying frames before anything leaves the hospital, and pulls performance analytics out of the footage afterwards.
Miscellaneous non-medical applications
Hospitals also run computer vision on jobs that have nothing to do with a diagnosis: counting what’s on the supply shelf, pulling handwritten intake forms into the record, and wayfinding for the visitor lost on the third floor.
The uncomfortable application is pointing the same cameras at your staff. Hand hygiene compliance is where most hospitals start, and the model is the easy half of it; the conversation with the union and the privacy officer is the rest.
What patients use computer vision for on their own phones
The consumer side runs on the camera already in someone’s pocket.

Everything on the consumer side runs on the camera already in someone’s pocket.
Preliminary research
Patients reach for these tools before they reach for a clinician, usually to work out whether the thing on their arm is worth an appointment.
Example: SkinVision assesses a photographed skin spot for cancer risk from an ordinary phone camera. It earned EU MDR Class IIa certification in August 2025, the tier that makes a vendor prove a medical purpose under notified-body review.
The cautionary example is Google’s. It demoed a skin-condition tool at I/O in May 2021, CE marked as a Class I device in the EU, never offered in the US, and Google said at the time it wasn’t intended to provide a diagnosis. Google’s health pages now list DermAssist as no longer in development, and the only skin feature it still ships is a Lens image-similarity search that Google describes as not a medical analysis.
Fitness & recovery therapy
Pose estimation puts a skeleton on a moving body from a single camera frame, which is enough to check whether someone is doing an exercise correctly. The same geometry drives elderly fall detection and ambient monitoring in a care setting.
Kaia Health’s Motion Coach uses the front-facing camera to find body landmarks during an exercise and corrects form out loud, with no wearable and no extra hardware. Amazon’s Halo held this slot until 2023, when Amazon ended support and the bands, the Rise unit and the app all stopped working.
Assistance to visually impaired people
People with vision impairment use computer vision to navigate unfamiliar places. Paired with a depth sensor such as Apple’s LiDAR, it builds a map of the space around someone and describes it out loud.
The OrCam MyEye 3 Pro is a 22.5-gram camera that clips magnetically onto a pair of glasses, reads text aloud on a voice command or a hand gesture, and identifies faces, products, colours and banknotes offline.
Related: vision is one slice of the picture. The full catalogue of generative AI healthcare use cases covers documentation, coding and patient communication too.
Main advantages of computer vision in healthcare
Four advantages come up over and over for computer vision in healthcare. Three of them hold up. Accuracy is where most articles overclaim.

The accuracy advantage is the one that needs its scope attached.
Clinicians get time back, which is the value-based care argument
Taking manual measurement and manual reading off a clinician’s desk gives them more of the appointment to spend on the patient, and lets them see more patients in the same day. Under value-based care, where reimbursement tracks outcomes, that reclaimed time is the part that shows up in the numbers.
Disease turns up earlier
A model holds every prior examination it was trained on and compares against all of them at once, which surfaces anomalies that are genuinely hard to see by eye. Find it early and treatment still works. That’s where most of the clinical value in this field sits.
Accuracy claims need their scope attached
Speed and accuracy both feed outcomes. The accuracy half comes with a scope the headline version drops.
In a 278-patient prospective trial published in Nature Medicine in 2020, a convolutional network reading stimulated Raman histology returned an intraoperative brain tumour diagnosis in under 150 seconds at 94.6% accuracy, against 93.9% for pathologists reading conventional slides. That is optical histology on fresh surgical tissue, running on equipment most hospitals don’t own. On MRI, where most people picture this happening, a 2026 review of 117 deep-learning studies found only 13% reported any external validation, with cross-dataset performance falling 2.5 to 15 points under domain shift.
Also read: Medical large language models: bridging the gap between technology and healthcare
Computer vision healthcare use cases by medical specialty
Computer vision in medicine splits by specialty. Each row below carries the named product and its regulatory status, so you can see which of these are cleared products and which are still research.
| Specialty | What the model does | Named example | Status |
|---|---|---|---|
| Radiology | Flags a large vessel occlusion on CT angiography and pages the stroke team | Viz LVO, Viz.ai | FDA De Novo DEN170073, February 2018 |
| Ophthalmology | Diabetic retinopathy screening from a fundus image, with no physician in the loop | LumineticsCore, Digital Diagnostics | FDA De Novo DEN180001, April 2018 |
| Cardiology | Coaches a non-expert operator through capturing a diagnostic-quality echo | Caption Guidance, GE HealthCare | FDA De Novo DEN190040, February 2020 |
| Digital pathology | Flags slide regions suspicious for prostate cancer for the pathologist | Paige Prostate, Tempus AI | FDA De Novo DEN200080, September 2021 |
| Surgery | Warns the surgeon when an instrument leaves the visible field on a live laparoscopic feed | Instrument Exit Point, Medtronic | FDA 510(k) K253984, June 2026 |
| Dermatology | Classifies a skin lesion as malignant or benign | Research and consumer apps only | No FDA-authorized autonomous classifier as of September 2026 |
| Musculoskeletal and rehab | Measures joint angles and range of motion from a patient’s phone video | Allheartz, built by Topflight | Commercial, as is the rest of the remote therapeutic monitoring category |
Dermatology is the gap between the claims and the clearances
Computer vision identifies benign lesions and tracks growth and colour change over time. The parity result everyone cites is real, and it’s 9 years old: Esteva and colleagues published it in Nature in 2017, training on 129,450 images and testing against 21 board-certified dermatologists, and Tschandl’s reader study followed in Lancet Oncology in 2019.
Both measured performance on curated, biopsy-proven image sets. External validation since has been harder on these models, and they underperform on Fitzpatrick V and VI skin, which is the population where a missed melanoma is already most likely. No autonomous image-based lesion classifier is FDA authorized in the US, for clinicians or consumers. If you’re building here, our guide to skin cancer detection app development covers what that constraint does to a product.
Radiology carries three quarters of every AI device on the market
X-rays, CT and MRI are where computer vision in healthcare is densest, and the FDA numbers show it: 1,230 of the 1,614 AI-enabled devices on the agency’s list went through the radiology panel. Radiology AI covers a spread of computer vision medical imaging jobs:
- disease classification
- image segmentation
- enhancement and reconstruction
- automated report generation
- temporal tracking across prior studies
Cardiology moved from reading the image to capturing it
Cardiac ultrasound and echocardiograms are the workload, and the shift here is upstream of interpretation. Caption Guidance was the first software FDA authorized to coach a clinician through acquiring a diagnostic-quality view in real time, which matters because a bad echo is the most common reason a good model gets nothing useful to read.
Surgery has one cleared product and a lot of demos
Surgical computer vision works on video: assessing technique, and putting contextual prompts in front of the surgeon as the procedure runs. Most of that is still research. The cleared exception arrived in June 2026, when Medtronic’s Instrument Exit Point went through 510(k) to warn a surgeon when an instrument drifts out of the visible field on a robotic procedure. Note how narrow that is: it tracks instruments while anatomy stays out of scope.
Pathology is where reader-to-reader variation is the problem being solved
Cancer detection, lung and breast especially, is the primary use of computer vision in pathology. Reading a slide is visual inspection, so it carries the subjectivity that goes with human visual perception, and it varies between pathologists.
The same models feed diagnosis and prognostication of outcomes and treatment response. Paige Prostate was the first AI pathology device FDA authorized, in September 2021, and it flags regions for the pathologist to look at.
Ophthalmology set the precedent everything else is measured against
Autonomous diabetic retinopathy screening is the only category where FDA has authorized software to produce a diagnosis with no clinician reading behind it. That started with LumineticsCore in 2018 and is now a competitive category: FDA’s diabetic retinopathy product code holds authorizations for Eyenuk’s EyeArt, AEYE Health’s AEYE-DS and, since July 2026, iHealthscreen’s iPredict-DR.
Models also read glaucoma, macular degeneration and childhood blindness from retinal images, and can predict non-ocular conditions such as anaemia and chronic kidney disease from the same photographs. Our eye health app development guide covers the capture and calibration problems that decide whether any of it works on a phone.
A computer vision feature is what turns your app into a medical device
Founders get this one wrong constantly, and it isn’t a close call. FDA’s Clinical Decision Support Software guidance, reissued in January 2026, carries a criterion that reads as if it were written for you: software “intended to acquire, process, or analyze a medical image” for a clinical purpose stays a device. The clinical-decision-support exemption that lets a lot of other health software off the hook is closed to you, however clearly you explain your reasoning and however carefully you show your work to the clinician.
The general wellness policy, also reissued in January 2026, is the other side of the line. FDA says it “does not intend to examine low risk general wellness products” whose claims stay on maintaining general health. A form-check feature in a fitness app lives there. Point the same model at a diagnosis and you’re filing with FDA.

The left branch is narrower than most founders assume.
Which pathway, and what it costs
Almost everything goes through 510(k). Of the 1,614 FDA-cleared AI devices on the agency’s list, 1,553 reached the market by matching a predicate, 40 through De Novo and 21 through PMA. Our published planning numbers for those routes:
| Pathway | When you use it | Planning range |
|---|---|---|
| 510(k) | A cleared predicate already exists that your device is substantially equivalent to | $50,000 to $300,000+ |
| De Novo | Nothing to match, and the risk is low enough for a new class | $150,000 to $500,000+ |
| PMA | Highest-risk devices, original safety and effectiveness evidence | $1 million to $5 million+ |
A De Novo creates the classification everyone after you clears into, which is exactly how Viz.ai’s stroke triage and LumineticsCore’s retinopathy screening went through. Both of those product codes now hold multiple competitors who arrived later by 510(k). Our guides to FDA clearance for health AI and to when an RPM device needs clearance walk the triage questions.
Your model ships frozen
FDA finalized its Predetermined Change Control Plan guidance in December 2024 and reissued it in August 2025. A PCCP lets you pre-clear a specified set of model updates inside the original submission, so a retrain that falls within the plan doesn’t need a new filing.
Read the scope carefully. A PCCP covers modifications you described and justified in advance. As FDA told its own Digital Health Advisory Committee, the agency has “yet to authorize any unlocked AI-enabled device.” Every AI device on the US market ships with its model frozen, and a change control plan governs when you’re allowed to change it. Budget for revalidation cycles.
Which tools you build on, and which ones can touch PHI
Almost every team builds on existing components when a software as a medical device project needs image recognition. Here’s what’s current in 2026, with the licence and the compliance status attached, because both of those decide things later that look free now.
- OpenCV (Apache-2.0, v5.0.0 since June 2026) is still the baseline library for everything around the model: preprocessing, geometric transforms, contour and edge work, and the classical detectors that still beat a network on simple shapes.
- PyTorch and torchvision (BSD-3-Clause) is the default once you’re training your own model.
- MediaPipe, now part of Google AI Edge (Apache-2.0), ships ready-made on-device vision tasks including pose landmark detection.
- YOLO from Ultralytics (YOLO26 current, YOLO11 the other production line) is the fast object detector everyone reaches for, and it’s AGPL-3.0. Ultralytics sells a separate enterprise licence precisely because AGPL’s obligations don’t fit inside a closed commercial product. Budget for that licence now, or find it during diligence.
- Hosted vision APIs, meaning Google Cloud Vision, Amazon Rekognition, and Azure Vision in Foundry Tools, which is what earlier versions of this article called Microsoft Computer Vision. Microsoft has flagged its Image Analysis 4.0 service as deprecated, retiring in September 2028.
- Multimodal model APIs for the metadata and drafting work covered above, with the caveats that come with it.
IBM Watson used to sit on this list and no longer belongs on it. IBM sold the Watson Health data and analytics assets to Francisco Partners in a deal that closed in June 2022, and they trade as Merative now, with the imaging products under the Merge name.
HIPAA eligibility is per service, and it’s yours to verify
A version of this advice circulates everywhere, and earlier versions of this page carried it too: the big clouds make their tools HIPAA compliant, so you don’t have to worry about it. It’s wrong, and it’s the kind of wrong that surfaces in an audit long after the demo went fine.
No cloud service is HIPAA compliant by virtue of who sells it. HHS is explicit that a provider creating, receiving, maintaining or transmitting ePHI is a business associate and has to be under a signed BAA, and that holds even when the provider only ever stores encrypted data it can’t read. Each vendor then scopes its BAA to a named list of eligible services and leaves correct configuration to you. Google’s own guidance tells customers to make sure they don’t use products that aren’t explicitly covered.
Teams get caught at the endpoint. Google Cloud Vision, Amazon Rekognition and Azure Vision all sit inside their vendors’ BAA scope. Google’s Gemini Developer API and AI Studio are absent from Google Cloud’s Covered Products list, which does name Generative AI on the Gemini Enterprise Agent Platform. Anthropic covers its Messages API while leaving out the Batch, Files, Skills, Code Execution and Computer Use APIs. OpenAI’s API becomes eligible once your organization is provisioned with Modified Retention.
Those lists change, this one included. Before a single image leaves your infrastructure, check the one that applies to you: Google Cloud’s Covered Products, the AWS HIPAA Eligible Services Reference, Microsoft’s Product Terms and Data Protection Addendum, and each model vendor’s own BAA page.

Check the vendor’s list before a single image leaves your infrastructure.
Two build decisions that change your compliance surface
Clinical images arrive as DICOM objects, and the header travelling with the pixels carries PHI of its own. NEMA maintains the standard and republishes it several times a year, PS3.1 2026d being current. De-identifying that header is a preprocessing step in your pipeline, and it has to run before the pixels reach your model.
Running inference on the device changes the picture entirely. Apple’s Core ML runs predictions, and optionally fine-tuning, on the person’s own handset. If the image never leaves the phone, no cloud provider becomes a business associate for that inference and there’s no BAA to negotiate for it. Edge AI costs you model size and update velocity, and buys you the cleanest data-flow story you can put in front of a security reviewer.
What a medical computer vision build actually costs
These come from projects we delivered, so treat them as planning anchors you can hold a vendor to.
| Scope | Range | What you get |
|---|---|---|
| MVP on an existing model | $25,000 to $80,000 | MediaPipe, a hosted vision API or an off-the-shelf detector, one platform, no custom training |
| Production build, AI-assisted | $25,000 to $97,000, $49,000 median | Our own healthcare projects built on pre-built HIPAA-ready modules |
| Fully custom build | $75,000 to $130,000+ | Custom architecture, EHR write-back, more than one surface |
| Custom model training, on top | $70,000 to $200,000+ | Data acquisition, labeling, validation, drift monitoring |
| FDA pathway, on top | $25,000 to $80,000 | Class I to II SaMD submission, adds 3 to 8 weeks |

The adders stack on top of whichever band you start from.
Four things move these numbers more than anything on a feature list.
Data labeling is the first, and it’s the one foundation models have genuinely changed. Annotating medical images means paying clinicians to draw boundaries, which is why the MedSAM figure earlier matters: model-assisted annotation cut two radiologists’ time by about 82% on the same task. Budget for annotation as a line item with a rate attached, because it behaves like a service contract.
The FDA pathway is the second, and it’s binary. A wellness feature ships when you ship it. A device feature adds a submission, a quality system and a clinical evidence question, which is the $25,000 to $80,000 adder above for the simpler classes and a great deal more past that.
Third: where inference runs. The tradeoff itself is in the tools section above; the budget consequence is that on-device work front-loads optimisation effort, while cloud inference spreads the cost across every scan you process.
Licensing is the fourth and the one that surprises people, because it arrives during diligence. AGPL-3.0 on the Ultralytics YOLO line is the common case. Our app development cost guide breaks down the rest of the budget, and machine learning as a feature runs $20,000 to $50,000 and up inside it.
What we learned building a computer vision RTM platform
Allheartz is a remote therapeutic monitoring platform for physiotherapy and sports care that we built end to end, computer vision included. A patient films a short exercise video on their own phone at home. The app runs pose detection on it, builds a skeletal model and extracts goniometric data, joint angles and range of motion, which reaches the clinician as measurements they can read in seconds.

Measurements reach the clinician in seconds.
The whole project turned on the machine learning. After R&D across the options we landed on TensorFlow and Google’s MoveNet pose detection model, running inference in the cloud to leave headroom for model updates. MVP v1 took 800 hours over 6 months.
Allheartz reports up to 50% fewer in-person visits and up to 80% less time spent on clerical work. If you’re building in this category, our RTM app development guide covers the CPT codes and the Class I device question, and rehabilitation app development covers the wider product around it.
If you’re weighing one of the computer vision healthcare applications in this guide, schedule a call to talk it through with our team.
Related Articles:
- How to create a physiotherapy app
- How to develop a skincare app
- Healthcare Cloud Computing Guide
- How to start a healthcare startup
- Healthcare App Development Guide
- How to create a telehealth app
- Python in Healthcare: Use Cases and Best Practices
[Reviewed September 2026]
Frequently Asked Questions
How is computer vision in healthcare used today?
Mostly for reading medical images: radiology accounts for 1,230 of the 1,614 AI-enabled devices on the FDA’s list. Past radiology it screens for diabetic retinopathy with no physician reading behind it, flags suspicious regions on pathology slides, totals surgical blood loss from photographs of sponges, coaches an operator through capturing a usable echo, and measures joint angles from a patient’s phone video at home.
What is the difference between computer vision and medical imaging AI?
Computer vision is the broader technique, any model that takes pixels as input. Medical imaging AI is the subset pointed at clinical images such as CT, MRI, X-ray, fundus photographs and pathology slides. A hospital camera checking hand hygiene is computer vision on its own; a model reading a chest X-ray is computer vision in the medical field and medical imaging AI at once, and only that second case puts you in front of the FDA.
Do computer vision medical apps need FDA clearance?
If the model analyzes a medical image for a clinical purpose, yes. FDA’s Clinical Decision Support Software guidance, reissued in January 2026, states that software intended to acquire, process or analyze a medical image stays a device, which closes the clinical-decision-support exemption to you. Features that keep their claims on general wellness sit outside that line. Of the AI devices FDA has authorized, 96% arrived through 510(k).
Can LLMs like GPT or Gemini analyze medical images?
Use them for the work around the image. In a 2024 Radiology study GPT-4V identified the imaging modality correctly 100% of the time, then returned 8% specificity for brain haemorrhage and an 86.5% false-positive rate on free-text reports. A February 2026 study in Diagnostics put GPT-5 Pro at 64.3% accuracy on fracture radiographs. No FDA-authorized device is identified as using a generative model. Routing, protocolling and drafting report text are the jobs these models do reliably today.
How much data do I need to train a medical CV model?
Less than you needed 3 years ago, with one caveat. Google reports its MedSigLIP encoder reaching strong chest X-ray classification from 512 or more labelled examples, and MedSAM’s own study cut expert annotation time by about 82% on abdominal tumour slices. Coverage matters more than raw volume: the standard failure in medical computer vision is a model that performs on the population it saw and degrades on the one it never did.
How much will it cost to build, let's say, a skin analyzing mobile app that would work in iPhone or Android apps?
Around $125,000 with an extensive image dataset for training, and an MVP fits within $80,000. The cost section above breaks the bands down by scope and lists what moves them.
How long does it take to develop a healthcare computer vision application?
4 to 5 months with ready-made computer vision libraries. Building it from scratch takes far longer, which is why most teams start from libraries. An FDA submission adds another 3 to 8 weeks for the simpler device classes.
Are there available tools for building mobile machine vision applications?
Yes. Apple’s Vision framework handles face and landmark detection, text, barcodes and general feature tracking, and Core ML runs the model on the device so images stay off the network. MediaPipe ships ready-made on-device vision tasks including pose landmark detection and runs on both iOS and Android.

