Podcast

QCast Episode 59: Clinical Data Quality

Written by Marketing Quanticate | Aug 14, 2026, 8:00:00 AM

Clinical data quality is often associated with clean databases, completed fields and resolved queries, but those measures only show part of the picture. For clinical trial data to be useful, they need to support the scientific question, analysis and decisions the study is designed to make. In this QCast episode, co-hosts Jullia and Tom examine how decisions made during study set-up, data collection and review can shape clinical data quality throughout the trial.

The discussion also looks at some of the practical pressure points behind poor-quality data. Electronic systems and automated transfers can reduce manual handling, but they also increase the need for clear mappings, definitions and ownership across systems. A technically successful transfer can still carry an incorrect unit or misunderstood source field downstream, while repeated queries may correct individual records without addressing the process that caused the problem.

🎧 Listen to the Episode:

 

 

Key Takeaways

Data quality starts before data cleaning

CRF design, collection guidance and clear data definitions influence the quality of the dataset long before formal cleaning begins. If sites interpret an endpoint differently or study instructions are unclear, a complete record may still be difficult to interpret consistently. Early design decisions can therefore reduce avoidable variation and downstream rework.

Automation does not remove the need for data oversight

Automated transfers can reduce repetitive manual errors, but they can also reproduce an incorrect mapping or definition consistently. Clear specifications, metadata and documentation help teams understand where a value came from, how it was transformed and whether it still represents the intended clinical information.

Quality review should focus on critical risks and patterns

Treating every data point as equally important can spread review effort too thinly. Greater attention should go to data that affect participant safety, key endpoints and major study decisions. Patterns such as repeated missing assessments or delayed adverse event entry can also reveal broader process issues that individual queries may not address.

Full Transcript

Jullia

Welcome to QCast, the show where biometric expertise meets data-driven dialogue. I’m Jullia.

Tom

I’m Tom, and in each episode, we dive into the methodologies, case studies, regulatory shifts, and industry trends shaping modern drug development.

Jullia

Whether you’re in biotech, pharma or life sciences, we’re here to bring you practical insights straight from a leading biometrics CRO. Let’s get started.

Tom

Today we’re discussing clinical data quality. Now when people hear this, I think the immediate assumption is mostly about making sure the data is complete, cleaning up errors and resolving the queries. From where we stand today, is that still a useful way of thinking about it?

Jullia

Somewhat, but only up to a point. Having clean and complete data is obviously important, but quality is really about whether the data is fit for what the trial needs it to do. Can it support the scientific question, the analysis, operational decisions and ultimately regulatory review?

Now that changes the conversation because a dataset can look very clean and still have weaknesses in how the data was really defined or collected.

Tom

Give me an example of that. What might look perfectly acceptable at first glance but still create a quality problem?

Jullia

So take an endpoint collected through an electronic form, for example. You could have every field completed and no outstanding queries, but if sites interpreted the question differently, or the wording didn't reflect what the protocol intended to measure, completeness hasn't solved the problem.

The same applies if a dosing change is recorded differently across systems. The entries themselves may be technically valid, but if their meaning isn't consistent, you have a problem when you try to interpret the data later.

Tom

So really, quality starts earlier than the usual image of data managers reviewing records during study conduct?

Jullia

Yes, much earlier. Decisions made during study set-up can have a considerable effect on what happens months later.

You need to decide what data is genuinely important, how it should be captured, what definitions sites will work to and what happens when that data moves between systems. Good CRF completion guidance and clear responsibilities across the study team can prevent a lot of avoidable variation before the first participant is even enrolled.

Tom

And I suppose the move to electronic systems hasn't removed those issues. It's just changed where they occur?

Jullia

Yes. While electronic data capture reduced some of the transcription problems associated with paper, trials now rely on much more complex data flows.

You might have EDC alongside ePRO, central laboratory data, imaging, electronic health records, wearables or other external sources. That puts more emphasis on integrations, mappings and definitions. The question is whether the right data arrived, in the right form, with their meaning intact.

Tom

There's a common assumption there that automation should improve quality because it removes manual handling. Is that always true?

Jullia

Not necessarily. Automation can reduce certain types of error, particularly repetitive manual transcription, but it can also reproduce a bad decision very efficiently.

Imagine laboratory data arriving automatically each night. The transfer works every time, there are no manual uploads and technically nothing fails. But if a unit has been mapped incorrectly or a source field has been interpreted differently from the receiving system, the automated process keeps passing that problem downstream.

This means somebody still has to understand what the transformation is actually doing, and that becomes especially important when several systems are involved. Teams need clear mapping specifications, metadata and documentation showing where a variable came from and what happened to it.

Tom

Does standardisation help with that at all? CDISC standards are an obvious example, but sometimes standardisation is treated mainly as something you do closer to submission.

Jullia

Well, it can be much more useful when considered earlier. If teams understand downstream requirements while designing the study and the collection tools, they can make more consistent decisions from the start.

But standardisation isn't only about putting data into a target structure. You still need clear source definitions, controlled terminology and consistent field meaning. As you’d imagine, a standard variable containing poorly defined source data doesn't suddenly become high-quality data.

Tom

Right. And once the study is running, this becomes less about preventing every possible issue and more about finding the ones that actually matter?

Jullia

Precisely. Trial teams can collect huge numbers of data points, and treating every one as equally critical isn't necessarily the best use of review effort.

Really, greater attention should go to data that affect participant safety, important endpoints or major study decisions. That's consistent with the wider quality-by-design and risk-proportionate principles used in clinical research. Identify what is critical and manage the risks around it.

Tom

So say a study team is halfway through recruitment. What might that kind of focused review actually look like?

Jullia

So they might look at whether key safety data is arriving on time, whether important endpoint fields are being completed consistently or whether particular sites show unusual patterns.

A trend review could reveal repeated missing assessments around a particular visit. Or perhaps adverse event records are regularly being entered several days late at one site. Those patterns can tell you more than checking isolated records one by one because they point towards a process issue that needs attention.

Tom

And presumably the response shouldn't automatically be another blanket round of queries?

Jullia

Correct. See, sometimes a query is appropriate. In other cases, the underlying issue could be training, unclear instructions or the way the collection tool has been designed.

If several sites make the same mistake, repeatedly querying individual records might fix the existing data but leave the cause untouched. That may mean clarifying the guidance or changing how teams are instructed to enter the information.

Tom

That brings us to quality control, because the term can become quite broad. What does quality control actually mean in day-to-day clinical data management?

Jullia

So really, it's the operational checking that happens as the study progresses. That can include programmed edit checks, listings review, reconciliation between systems, query management and review of unusual values.

The aim is to identify data that are missing, implausible or inconsistent while there's still a reasonable opportunity to investigate and correct them.

Tom

And how is that different from quality assurance? They're often used almost interchangeably in conversation.

Jullia

Well, they're closely related, but the emphasis is quite different. Quality control is generally concerned with finding issues in the data or study processes during conduct.

Quality assurance is more about whether the right processes and controls were established in the first place. So, for example, having clear specifications, defined responsibilities and an appropriate review process supports quality assurance. Running a check that identifies an inconsistent visit date is quality control.

Tom

And when a process changes during the study, how much does training matter?

Jullia

Training only helps when it's attached to a clear process. So generic onboarding followed by repeated reminders won't compensate for ambiguous study instructions.

It only really becomes more valuable when something changes. If there's a protocol amendment affecting a visit schedule, for example, the relevant teams should understand exactly what changed in data collection and review. The same applies when an eCRF or another study document is amended.

And you’d be right in thinking this sounds as much like an ownership issue as a technical one. Somebody has to know who is responsible when a problem crosses systems or functions.

See, modern data quality depends on coordination. If laboratory data don't reconcile with the EDC, or an external data transfer arrives late, there needs to be a clear route for resolving it. Technology can flag the discrepancy, but it can't decide responsibility or resolve an unclear definition between teams, which is why governance still matters.

Tom

Now if someone listening is reviewing their own study, what's a useful way to test whether their approach to data quality is actually working?

Jullia

I'd start with the data that matters most to the study and ask whether everyone understands what that data means, how it’s collected and how it’s reviewed.

Then follow a few of those data points through the lifecycle. Take something important such as a primary endpoint assessment or a safety laboratory value. Ask yourself, where is it entered? Does it move between systems? Is it transformed? What checks apply, and who acts if something looks wrong? That quick exercise tends to expose gaps quite quickly.

Tom

So really, you're following the data rather than starting with a checklist of every quality activity the organisation performs?

Jullia

Exactly. A long list of controls doesn't necessarily tell you whether the important risks are controlled. So if we were to reduce today's discussion to a few takeaways, I'd keep it fairly simple. Define quality according to how the data will be used. Build quality into study design rather than relying on cleaning later. And where data moves across systems, make sure its meaning and source remain clear.

With that, we’ve come to the end of today’s episode on clinical data quality. If you found this discussion useful, don’t forget to subscribe to QCast so you never miss an episode and share it with a colleague. And if you’d like to learn more about how Quanticate supports data-driven solutions in clinical trials, head to our website or get in touch.

Tom

Thanks for tuning in, and we’ll see you in the next episode.

About QCast

QCast by Quanticate is the podcast for biotech, pharma, and life science leaders looking to deepen their understanding of biometrics and modern drug development. Join co-hosts Tom and Jullia as they explore methodologies, case studies, regulatory shifts, and industry trends shaping the future of clinical research. Where biometric expertise meets data-driven dialogue, QCast delivers practical insights and thought leadership to inform your next breakthrough.

Subscribe to QCast on Apple Podcasts or Spotify to never miss an episode.