Skip to main content

Scientific workflows for primary care database studies

A scientific workflow is a method in computer science for formulating abstract descriptions of analytical processes. This allows automation and reuse of many of the tasks in analysing large, complex data sets. A recent article in the journal Statistical Methods in Medical Research discussed the use of these scientific workflows for the analysis of data from large primary care databases.

Routinely collected primary care data in electronic repositories are a promising source of data for audits, quality improvement, health service planning, epidemiological studies and research. However, a number of challenges have been noted about working with these data sets. In the paper, we discuss these issues and describe how we used scientific workflows to analyse data from one large primary care database (GPRD). Some of the steps in the analysis of data from the GPRD our shown in the figure below.

Comments

Popular posts from this blog

What is the difference between primordial prevention and primary prevention?

Primordial prevention and primary prevention are both crucial strategies for promoting health, but they operate at different levels. Primordial prevention aims to address the root causes of health problems and improve the wider determinants of health. It focuses on preventing the emergence of risk factors in the first place by tackling the underlying social, economic, and environmental determinants of health. This involves broad, population-wide interventions such as: Policies that promote healthy food choices: Think about initiatives like taxing sugary drinks to discourage unhealthy consumption, or providing subsidies for fruits and vegetables to make them more accessible. Urban planning that prioritises well-being: This could include creating walkable neighborhoods with safe cycling routes, ensuring access to green spaces for recreation and relaxation, and designing communities that foster social connections. Social programs that address inequality: Initiatives aimed at reducing pov...

What makes a good doctor – and who gets to decide?

What Makes a Good Doctor? This is the question that Waseem Jerjes and I explore in the Journal of the Royal Society of Medicine . It is a key question that underpins the architecture of medical education, clinical practice, regulation, and professional identity. It cannot be answered by regulators, educators, or employers in isolation. It must be answered together – by doctors and patients – revisited throughout a career, and adapted as society and the profession change. Without that shared reflection, the danger is not simply disillusionment, but the erosion of the moral foundations of clinical work. As we enter an era when diagnosis will increasingly involve artificial intelligence and when performance metrics reward volume over value, reclaiming this question as a professional one is imperative. The integrity of our institutions – and of the practitioners within them – depends on reimagining excellence in inclusive, relational terms. A good doctor is not a flawless technician or a f...

Relevance Over Recall: Rethinking How AI Uses Clinical Data

Our article in the Journal of the Royal Society of Medicine argues that safe and effective AI in healthcare must incorporate mechanisms that emulate human judgement - down-weighting old, inaccurate or superseded information and prioritising what is recent, clinically relevant and reaffirmed - so that AI supports, rather than disrupts, high-quality patient care.  Clinicians constantly revise, reinterpret and filter past information so that only what is relevant, accurate and timely shapes present-day management decisions; medical records function as dynamic “working tools” rather than fixed archives. By contrast, many AI systems lack this capacity for selective forgetting and often treat all historical data as equally meaningful.  This can lead to outdated or low-confidence diagnoses being repeatedly resurfaced, persistent labels influencing clinical expectations, and irrelevant, long-resolved events cluttering summaries and decision-support outputs. Such indiscriminate recall...