February 5, 2013

Complexity Analysis of MU Stage 2 Eligible Professional Clinical Quality Measures

I recently wrote about the initial release of the MITRE open source Kamira project to assess the quality of Clinical Quality Measures (CQMs).  The initial focus of the Kamira project has been the Cyclomatic Complexity analysis of the CQM logic.

The Kamira project has listened to my evangelizing the use of Kiviat charts as the visualization technique used with CQMs.  In addition to being fairly useful for understanding the CQM complexity results, I have found that the Kiviat visualizations also have a "sexiness" factor that I have been leveraging to continue the research.  See the Kamira Cyclomatic Complexity dashboard design below:

Kamira complexity dashboard design
Kamira complexity dashboard design
What we did with the Kiviat visualization for the complexity results is leverage the five attributes that could always be used for parts of the CQM algorithmic logic:
  • Initial Patient Population (IPP)
  • Denominator 
  • Numerator 
  • Exclusion 
  • Exception
We then associated ranges for the complexity metrics based on the Carnegie Mellon University paper that I identified a few months ago.  These ranges are:
  • 1-10: Very Simple, Low Risk (green)
  • 11-20: Nominal, Moderate Risk (yellow)
  • 21-50: Complex, High Risk (orange)
  • >50: Untestable, Extreme Risk (red)
Since we wanted to highlight the worst/highest component of the CQM, we opted to color in the area of the five metrics with the worst/highest section of the algorithmic logic.

Because some of the MU Stage 2 CQM logical sections far exceeded CMU's threshold of 50 for "Untestable, Extreme Risk" we found that there were some CQMs that had logical sections that were "off the scale" when it came to Cyclomatic Complexity.  The worst violator is "NQF 0038: Childhood Immunizations" which is actually part of the core set of MU Stage 2 for Eligible Professional CQMs.

Since we couldn't use a linear scale for those CQMs, so we capped the scale at 60, and switched the presentation of the results around by providing a red background around the Cyclomatic Complexity value.  Normally, there isn't a background and the results are just black text.  The idea here was to try and call out that there was something fundamentally wrong with complexity values that are so high.

See the 6 most complex CQMs for MU Stage 2 Eligible Professionals (EP) below:

6 Most Complex CQMs for MU Stage 2 Eligible Professionals
6 Most Complex CQMs for MU Stage 2 Eligible Professionals
From a software engineer's perspective, if you have to manually implement these CQMs in software code, you are going to be very hard pressed to test and validate that any system implementing this logic accurately.  Ie. you have a challenging job ahead of you to test that any implementation of these 6 CQMs are in fact accurate.  This also means that there's a higher probability of having a lurking software bug in your implementation.

Also, for healthcare providers who need to understand these CQMs in order to improve the quality of care that they are providing to their patients, they will also have a challenge.  Providers needing to maintain a mental model of the CQMs and understand what actions should be taken for their patients will likely face challenges.

Ideally, the complexity of the CQMs should never even get this high.  What I would want to see is some consideration by measure stewards on Cyclomatic Complexity during the development process.  Another consideration for the policy leaders would be refusal to accept any CQM submitted that had a Cyclomatic Complexity value that exceeded a threshold.

The Kamira project's work has initially focused on the Meaningful Use Stage 2 Eligible Professional CQMs.  The reason that these were initially selected is because we were easily able to instrument the JavaScript code that is implemented in the popHealth and Cypress project's quality measure engine.  However, the work can be expanded to include the MU Stage 2 Eligible Hospital (EH) CQMs when those are implemented in popHealth and Cypress, currently planned for late April 2013.  Unfortunately, I expect that the EH CQMs will actually be more complex than the EP CQMs.

Additionally, I have plans to expand the work into financial analysis and feasibility of implementing CQMs, but for now the lowest hanging fruit was the complexity analysis.

The full visualization of the Meaningful Use Stage 2 CQM Complexity Data is available here, and the detailed results in a JSON file is available here.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. © Rob McCready, 2013.
Creative Commons License

February 3, 2013

Complexity and Certification of MU Stage 1 Eligible Professional Clinical Quality Measures

While working in the Clinical Quality Measure space on the two open source projects popHealth and Cypress, I have observed trends in the adoption of EHR vendors of various Clinical Quality Measures (CQMs).  For a little background on the role of CQMs in the Meaningful Use program, read the beginning of the entry that I wrote about Applying Kiviat Visualization to Meaningful Use Clinical Quality Measures. 

For Meaningful Use Stage 1, there are minimal requirements by EHR vendors to support the 6 Core and Core Alternate CQMs, and any 3 of the remaining 38 Meaningful Use Stage 1 CMS.  This allows for some malleability by the commercial EHR vendors to select CQMs based on either their ability to implement CQM logic in their product or based on customer demand for specific CQMs.

The Office of the National Coordinator for Health Information Technology (ONC) hosts the Certified HealthIT Product List (CHPL pronounced "CHaPeL") service.  You can view the list of certified products through the CHPL web interface.  Compiling the results for the EHR products against various Meaningful Use Stage 1 CQMs, there are some interesting results:

Meaningful Use Stage 1 Ambulatory Clinical Quality Measures:
Adoption by EHR vendors from data collected via the CHPL service

In addition to the work I am leading via Cypress, I am also leading a research project to assess the quality of Clinical Quality Measures called "Kamira".  The Kamira project can provide metrics on the quality of CQMs.  For early 2013, this has included automated Cyclomatic Complexity calculation of the CQM algorithmic logic based off of the JavaScript code that the Cypress and popHealth projects use to calculate the CQM results.

If you then compare the results of the CQMs that were tested and certified by EHR vendors against the complexity score of the CQMs, you can see a weak correlation between the two.  You can download the full file with the results here.

Meaningful Use Stage 1 Ambulatory Clinical Quality Measures:
Adoption by EHR vendors from data collected via the CHPL service
compared against Cyclomatic Complexity Analysis of the CMQ logic
It's worth noting that the correlation here is weak, but there does appear to be a trend toward vendors opting to implement the less complex CQMs in their products when they have some latitude to choose.

FYI, the ranges for the CQM complexity (the colored diamonds) are:
  • 1-10 Very Simple/Low Risk (green)
  • 11-20 Nominal/Moderate Risk (yellow)
  • 21-50 Complex/High Risk (orange)
  • >50 Untestable/Extreme Risk (red)
These metrics are somewhat arbitrary, but I picked ranges for CQM complexity from a Carnegie Mellon paper that had a good number of citations, so I think that the values and thresholds are fairly defensible.

Where this work might have a few vulnerabilities is that I am 100% certain that EHR vendors do not use complexity as their only consideration when selecting CQMs to implement in their products.  For instance some of the red, highly complex CQMs which were in the middle when it came to adoption by EHR vendors are cardiac CQMs.  From my perspective, it's a safe assumption that some of these EHR vendors were going to bite bullet and implement the cardiac CQMs regardless of the complexity associated with them because there is more demand in the marketplace from providers that need the cardiac CQM results, vs. say the behavioral health CQMs.  However, I think that CQM developers need to start tracking complexity of CQMs as they are developed for MU Stage 3 or beyond.

Lastly, the Kamira project just launched last week.  I plan on posting the MU Stage 2 complexity results for the Eligible Professional CQMs in the coming weeks.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. © Rob McCready, 2013.
Creative Commons License

November 25, 2012

Longitudinal Visualization Techniques for Clinical Quality Measures

I recently wrote about Applying Kiviat Visualization to Meaningful Use Clinical Quality Measures.  One additional item that is a limitation to using Kiviat diagrams is the visualization of longitudinal trending of metrics over time.  This same limitation exists with Clinical Quality Measures (CQMs).

I am a big fan of Edward Tufte, and his evangelism on the use of sparklines.  Below is an interpretation of how that same sparkline visualization could be applied to a family of diabetic CQMs  from the Meaningful Use program, that is a modified version of Juhan Sonin's HealthCard:


This technique does make some assumptions.  It appears to look decent with two years of notional data.  However, I think that changes in CQM results will probably be needed over decades.  Also, I am not sure if someone like Tufte would violently object to the introduction of date under the illustration.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. © Rob McCready, 2012.
Creative Commons License

October 7, 2012

The QRDA Category 1 XML Standard


The Quality Reporting Document Architecture (QRDA) Category 1 standard is a new XML standard designed for communicating patient-level clinical data that will be used to calculate Clinical Quality Measures (CQMs). This is an HL7 XML standard that is part of the Clinical Document Architecture  (CDA), which is an overall framework for expressing healthcare data in XML that is also stewarded by HL7.  

The QRDA Category 1 XML standard was created because systems needing to calculate Clinical Quality Measures (CMSs) could not guarantee that patient-level data would appear in the existing continuity of care standards such as the HITSP C32, HL7 Consolidated CDA (CCDA), and the ASTM Continuity of Care Record (CCR).  The QRDA Category 1 standard has been designed so that all of the clinical information needed for each specific CQM will need to be expressed within the QRDA Category 1 XML document. 

The QRDA Category 1 is a clinical document representing a single patient. It defines sections containing measure information and patient data that meets criteria for a measure. It also contains information on the reporting period. The markup for the patient data is the same as a CCDA document.  

In a nutshell, the QRDA Category 1 XML standard will include the templates needed for each attribute in each CQM that a patient could provide clinical data to be used to calculate any number of CQMs.  Additionally, more than one CQM can be included in the request for the clinical data associated with one patient.  

If you view the illustration below, you can see that an Electronic Health Record (EHR) system may be requested to generate QRDA Category 1 XML for patients to be used to drive the calculation of the Diabetes Blood Pressure Management CQM AND the Diabetes LDL Test CQM.  In that case, all clinical data that could be used to drive the calculation of either of those CQMs must appear in the QRDA Category 1 for each patient in the EHR.  Further, additional clinical information associated with each patient is not required, and will result in warnings (but not errors) in the QRDA Category 1 for each patient.

c32/ccda vs qrda category 1
The C32/CCDA vs QRDA Category 1
Only relevant data is expressed in the QRDA Category 1
based on the need for specific CQMs
I personally feel that the QRDA Category 1 reflects the failures of the CCR, C32, and CCDA standards.  It is difficult to explain why a single standard for expressing patient-level data that is interoperable and can provide sufficient fidelity of data for calculating CQMs does still not exist.  It's 2012.

Where problems are identified with a CQM when the CCR... the C32... the CCDA do not provide sufficient information for calculating MU CQMs, CQMs could be adjusted to be simpler, and conform to data that must exist in the CCR... the C32... the CCDA.  This lack of a good continuity of care standard for a single patient record is what I feel is the root cause of all of this work.  

Knowing that the QRDA Category 1 isn't going away anytime soon, I have found that the use of the QRDA Category 1 XML standard has been difficult due to lack of validation tests and example QRDA Category 1 XML files.  However, I suspect that these two shortcomings will be addressed in time.

The objective of using the QRDA Category 1 XML standard as a mechanism to present clinical quality data associated with a patient should unambiguously define clinical attributes so that there’s no confusion about what is expected of an EHR system's ability to export in XML artifacts.  Time will tell if this objective is realized in EHR systems... I remain skeptical.  I think we will know a lot more by 2013, when this standard is officially introduced into the Meaningful Use Stage 2 program.


This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. © Rob McCready, 2012.
Creative Commons License

September 13, 2012

Review: CardioChek PA Blood Meter

I have been working for the Office of the National Coordinator for Health Information Technology (ONC), on the popHealth and Cypress projects for several years now.  Both projects are based around Clinical Quality Measures (CQMs).

By education, I am a physicist and engineer, and not a clinician.  However, I have repeatedly noticed the importance of lipid profiles within the context of the logic of several of the Meaningful Use CQMs.  Further, in the very recent past, these metrics had been considerably high for me; a male in my late 30's.  Between my work looking at CQMs emphasizing the importance of just measuring lipid profiles, and my own personal warning signs associated with this health metric, I thought it would be good to collect some "hands on" experiences with the collection of the data for this metric.

When my department at MITRE had some additional overhead resources available at the end of our fiscal year, I purchased a portable blood testing device that could provide me with my own lipoid profile information (total cholesterol, HDL cholesterol, LDL cholesterol, and triglycerides).

I eventually picked the CardioChek PA blood meter device.  Interestingly, this was not the first CardioChek blood device that I purchased.  I originally found the consumer CardioChek (no "PA") device on Amazon.com.  The consumer device lists at about $125.  That seemed reasonable, but the device requires three separate strips to measure total cholesterol, HDL cholesterol, and triglycerides (you can calculate the LDL cholesterol from those three).  While that may not sound bad, I quickly found out that it involved giving myself at least two, sometimes three, pricks with lancets to draw enough blood for the three separate tests.

Additionally, I purchased this device for my department at work, and felt that an experience getting three pokes with a needle might not go over well with my colleagues, resulting in some future retribution via office pranks.

To solve the problem of collecting a full lipoid panel from one drop of blood, I purchased these PTS panels that appeared nice in the sense that they have the ability to drive the calculation of multiple readings from a single sample of blood.

Each box comes with one lipid panel MEMo chip that can be inserted
into a CardioChek PA device and 15 single use lipid profile panels

Unfortunately, I quickly discovered that these really nice 3-in-1 lipid profile panels are not compatible with the $125 consumer CardioChek device.  To support multiple readings from a single drop of blood would require the more expense CardioChek PA device, that runs at close to $700.

Being the end of the fiscal year, the additional money was a little easy to come by.  I went back to our finance staff and purchased the more expensive, clinical-grade device.  I also picked up some supporting medical equipment like gloves, lancets, pipettes, and band aids.

CardioChek PA blood testing device
CardioChek PA blood testing device, several tests
with some additional medical equipment

To take your lipid profile, you need one lipid panel test chip, called a MEMo Chip by the manufacturer, and one lipid panel test strip.  The MEMo chip contains lot-specific calibration and other information needed to properly perform testing.  The lot-specific information is presumably associated with the 15 test panels that come with the package.  I would not recommend mixing and matching test panels with different MEMo chips, because of this prior calibration by the manufacturer.

There is also guidance around always storing the unused test panels at a temperature between 68-80˚F.  I could see this temperature requirement as a challenge for some home/consumer users.  Lastly, there is an expiration date on the test panels.  For all the tests I have, the expiration date, is less than one year from now, a little under 8 months of viable shelf time.

You can see the relative size of the test panel and MEMo chip in the picture below.  You only need one strip for a test.  I just flipped one test strip over to show the single wide channel where you deposit your blood, and the three openings for the device sensor to read the total cholesterol, HDL cholesterol, and triglycerides.

lipoid panel test chip, with two lipoid panel test strips
Lipoid panel test chip and two lipoid panel test strips

It appears that the CardioChek PA device has the ability to independently test up to four different metrics from the single sample of blood.  While I haven't found any tests that use a full four metrics from one sample, I am still happy that the manufacturer (presumably) recognized this need for multiple tests from a single sample.

CardioChek PA sensor

Collecting and depositing your sample for the device is relatively easy for an individual, non-clinician.  I suggest you get 2 paper towels, and a small band aid before you get started.  If you are unfamiliar with using lancets, they are small and cheap medical implements used for capillary blood sampling (no vein or arties). A lancet includes a spring-loaded needle.  When used, it pops out and makes a very small puncture in your skin, allowing a few drops of blood to appear on your skin over the next ~10 seconds.  They are single use and disposable.

You can see how they are fairly straight-forward to use for the blood sampling that you can then collect with pipettes.

Lancet
About one drop of blood after lancet met my finger
with small pipette in background

After my first two attempts to use the device for a full lipid profile, the device would eventually display "TEST ERROR" on the screen.  Needless to say, I was disappointed to think that I had put over $1K into this exercise, and had no data to show.  As it turns out, my finger was not providing enough blood to generate the full lipid profile.  There is some documentation provided with the device that makes this association from the "TEST ERROR" message.  However, I just can't understand why the manufacturer didn't make this more intuitive for users.

On my third attempt, when I applied a liberal amount of blood on the sample channel, the device worked fine.  I feel that the accuracy of the device is very high.  The readings it provided were all within 10% of measurements that I had collected from my Primary Care Provider (PCP) the previous week.  For me, the time between depositing your blood on the strip until the results are displayed ranged from 40 to around 60 seconds on three test runs.

This problem that resulted in multiple "pokes" with the lancet makes for an amusing story at my expense.  However, I can't emphasize enough; the "TEST ERROR" message really means "Not enough blood".  This is an opportunity for improvement in this product for first time users.

On this topic of human-computer interaction, I was also disappointed with the CardioChek PA user interface.  For $700, the interface feels like bleeding-edge late 1980s technology.

CardioChek PA user interface

For a $700 device, I would think that the manufacturer could easily upgrade the resolution and include color at a modest increase in manufacturing costs.  Ideally, this would include some longitudinal data on changes associated with the data.  See my suggested illustration below, again based on Juhan Sonin's designs for a HealthCard for the patient (consumer).

My updated lipid profile with sparkline visualizations
On the positive side, the CardioCheck PA is worth purchasing if you think you will be frequently taking your lipid profile at home.  It appears very accurate on the measurement results.  Once you learn the interface and how to correctly provide enough of a blood sample, the device works well.

My biggest issue with this product is the variance in the cost of the CardioChek PA device at $700, versus a slightly more limited consumer CardioChek device at $125.  I feel that this price point makes the CardioChek cost-prohibitive for most consumers (patients).

I would grade the CardioChek PA device a B- for the purposes of home users.

Somewhat related to where this may go in 5-10 years, I was able to identify an amazing illustration, showing various ranges for numerous blood metrics:

Reference ranges for blood tests

Knowing that a single drop of blood could theoretically yield all these metrics makes for some interesting ideas about how the consumer could have access to these metrics on a daily basis, at their home.

Another interesting opportunity for this market would be to introduce a blood sensor without the embedded interface that could communicate with an iPhone, similar to how the Withings BP Cuff works.  I have that Withings device at home, and plan on developing a review of that device later.

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. © Rob McCready, 2012.
Creative Commons License