Showing posts with label multicriteria decision making. Show all posts
Showing posts with label multicriteria decision making. Show all posts

Monday, July 26, 2021

Reducing Noise and Improving Decision Making

 

Cover page image: https://www.littlebrownspark.com/

Noise: A Flaw in Human Judgment. By Daniel Kahneman, Olivier Sibony, and Cass R. Sunstein

Kahneman, Sibony, and Sunstein have written a book that is both valuable and frustrating.

This book presents multiple ideas related to human judgment and decision making: (1) a review of studies that have described the variability in judgments in many domains, (2) approaches for reducing that variability, (3) an approach for making decisions when there are multiple factors that should be considered, and (4) an appeal for better procedures in the legal system.  According to the authors, the book offers an understanding of the psychological foundations of disparities in judgments, which are here classified as noise and bias.

The book’s strengths include its review of the literature on the variability of judgments and its distinction between noise and bias.  It presents examples of judgments from many domains (including medicine, business, and the legal system).  It strongly supports systematic decision-making processes (a topic that is important to me) and emphasizes the importance of accurate judgments.  It acknowledges the difficulties of reducing noise.  Finally, its notes provide references to original studies that provide context and details to the book’s discussion.

The book describes a range of best practices for judgment and decision making: employing persons who are better at making judgments, aggregating multiple judgments, using judgment guidelines, using a shared scale grounded in an outside view, and structuring complex decisions.  The first three items are meant to reduce judgment errors due to noise and bias.  The last item is a practical multi-criteria decision-making process that (a) decomposes the decision into a set of assessments, (b) collects information about each assessment independently, and (c) presents this evidence to the decision-maker(s), who may use intuition to synthesize this information and select an alternative.  In an appendix, the authors use their recommendations, for the book provides a checklist (a guideline) for evaluating a decision-making process.

When discussing ratings, the book wisely recommends that “performance rating scales must be anchored on descriptors that are sufficiently specific to be interpreted consistently.”  Scales with undefined terms such as “poor” and “good” and “excellent” should be discarded unless they are well-understood in the group of persons who are using them as a common language.

Unfortunately, two weaknesses (one minor, one major) frustrated me.  The first is the fact that the text contains no superscripts, citations, or other marks that indicate the notes that are available in the Notes section at the end of the book.  In the Notes section, each note has only a page number and a brief quote that suggest the text to which the note applies.  This unreasonable scheme reduces the value of the notes' many citations and explanations by making them harder to find and use.

The second weakness is more significant.  The book does not distinguish between judgment and decision making.  It tends to treat them as the same thing.  Indeed, a note explains that the authors “regard decisions as a special case of judgment” (page 403). 

For example, the book discusses the judgments that an insurance company’s employees make.  One example is a claims adjuster’s estimate of the cost of a future claim, which is indeed a judgment.  The other example is the premium that an underwriter quotes, which is a decision, not a mere judgment.  It is based on numerous judgments, of course, but the underwriter chooses the premium amount.  The book states that making a judgment is similar to measuring something, which is appropriate, but then goes on to say that the premium in the underwriter’s quote is also a judgment, which is not appropriate, because it is the result of a decision, not a measurement.

Elsewhere, the book claims that “an evaluative judgment determines the choice of an acceptable safety margin” for an elevator design (page 67).  This is inappropriate, however, for choosing the safety margin is a decision, not a measurement.

The book states that the process of judgment involves considering the given information, engaging in computation, consulting one’s intuition, and generating a judgment; that is, judgment is “an operation that assigns a value on a scale to a subjective impression (or to an aspect of an impression)” (page 176).  This is not the same as decision making.  Decision making is a more comprehensive process that defines relevant objectives, identifies (or develops) alternatives, evaluates the alternatives, selects one, and implements it.  In this process, judgment is an activity that may be used to evaluate the alternatives.  The book provides a relevant example that shows the distinction: job candidates get ratings, but only one gets hired.  The ratings are judgments, but choosing and hiring someone is a decision.

Those seeking to improve decision making in their organizations will find many useful suggestions in this book, but they should keep in mind that decision making is a process, not a judgment.

 

Wednesday, May 6, 2020

Scary Spider Chart

Spider chart.  Creator: redacted.
The spider chart shown in this post came from a dissertation that studied multi-agent systems. A spider chart is also known as a radar chart, and writers use it to graph multivariate data.  In a typical example, each alternative has a polygon that connects the points that represent its performance on multiple measures, with one point on one spoke for each measure.  It can be effective to show how one alternative performs well on one or more measures and another alternative performs well on another set of measures.  In that case, the polygons for these alternatives will have different shapes.
Conversely, the polygons for two alternatives will be very similar if the performance of those alternatives is similar across the different measures.

Unfortunately, the spider chart shown here has reversed the typical use.  There are two alternatives (adaptive risk and fixed risk) that were tested under three scenarios (low, medium, and high workload).  In this chart, there are six spokes, one for each combination; a typical chart would have six polygons.  Instead, there are three polygons, one for each measure: profit, completed percentage, and failed percentage.  The last two measures always add to 100%, so one of them is redundant.

Thus, there are only twelve useful data points (two measures for six combinations). The data-to-ink ratio is very low.  Given the small amount of data, a simple table may be the best way to convey this information.

The purpose of this spider chart is to show how the two alternatives compare on the performance measures, but this chart makes that comparison very difficult. A reader normally relies upon the slope of a curve (either positive or negative) to determine how performance is changing, but that will not work here because the different scenarios have different orientations and one performance measure is the complement of the other. 

Because there are effectively two performance measures, a two-dimensional scatter plot (with appropriate labels) would have been appropriate.  The second chart (which I created using the same data) is a possibility; this makes the change from fixed risk to adaptive risk more clear, but it still has a low data-to-ink ratio.
Two-dimensional scatter plot.  Creator: Jeffrey W. Herrmann.



Monday, November 30, 2015

How to plan a test

Testing generates information that can be used to make a decision.  Testing can occur at any stage in the development of a product or system; it can test a specific attribute or overall performance; it can test a component, a subsystem, a system, or a system-of-systems.

Planning a test requires making crucial decisions: Which item to test?  Which test to perform?  How many tests to perform?

The test plan determines the value of the information gathered and the time and cost of testing.  In general the key tradeoff is that gathering more valuable information requires more time and cost.

My students and I have developed some techniques for making test planning decisions.

Which facility to use?  In some cases, a system (such as a military vehicle) needs to be used in an operational environment, but there are no existing test facilities like that environment.  Instead of building a new test facility, one could use a combination of existing facilities to replicate the new operational environment.  We developed an optimization model to find the test plan that used the best combination of existing facilities examples.  For military vehicle applications, it specified the time and number of miles that the test vehicles should run on each existing test facility.  Reference: http://www.isr.umd.edu/~jwh2/papers/IEST.pdf

Which configuration to test?  A system-of-systems (SoS) consists of relatively independent systems.  For example, a missile defense systems has control stations, radar locations, and rocket launchers.  If the reliability of the SoS is unknown because the reliability of the systems is unknown, then testing is needed, but testing a full-scale configuration is expensive.  We developed a simulation technique to predict the results of tests with smaller configurations under different scenarios and estimate the expected error of the tests.  With this information, the test planning decision-makers could evaluate the tradeoffs of cost and expected error and determine the best test configuration.  Reference: http://www.isr.umd.edu/~jwh2/papers/Tamburello-Herrmann-JRR-2015.pdf

Which attribute to measure (test)?  In a multiattribute decision making situation, measurements of the attributes are valuable for knowing which alternative is best.  If these measurements have error and the total budget for measurements is limited, then it is crucial to measure the attributes in a way that provides the most valuable information and increases the likelihood of selecting the truly best alternative.  We developed and tested rules for determining which attributes should be measured how many times and showed that better rules can significantly increase this likelihood.  Reference: http://www.isr.umd.edu/~jwh2/papers/Leber-Herrmann-ISERC-2015.pdf

Which test to perform?  Demonstrating the reliability of a complicated system (such as a liquid rocket engine) requires testing the system and its components and subsystems.  These tests and the associated hardware are expensive and require time at scarce test facilities.  We developed a multi-objective test plan optimization approach to determine the best test plan that meets the demonstrated reliability.  Reference: http://www.isr.umd.edu/~jwh2/papers/Strunz-Herrmann-CEAS-Space-2011.pdf









Tuesday, March 31, 2015

California's offshore oil and gas platforms

An article by Max Henrion in the February 2015 issue of OR/MS Today described the decision analysis used to determine the best option for decommissioning 27 offshore oil and gas platforms off the coast of Southern California.  A complete report on the analysis can be found here.

The article illustrates two of the three critical perspectives on decision making: (1) the problem-solving perspective (what do with the platforms) and (2) the decision-making process perspective (how to make the decision).

From the problem-solving perspective, the decision is actually 27 decisions, one for each platform.  The article lists multiple options in three categories: complete removal, partial removal, and leave in place for reuse.  Within each category were multiple alternatives. 

The attributes used to evaluate the alternatives were costs, air quality, water quality, impacts on marine mammals, impacts on birds, impacts on the benthic zone, fish production, ocean access, and compliance with lease terms. 

Based on the stakeholders' preferences, the analysts created a multi-attribute model.  For each platform, each decommissioning alternative was given a score (on a 0 to 100 scale) for each attribute, and the scores were combined using a weighted sum.  The alternative with the best total score was identified as the best for that platform.

From the decision-making process perspective, the process was an analytic-deliberative one, and the analysis involved many traditional tools, including influence diagrams, decision trees, sensitivity analysis, and swing weighting.

A multidisciplinary analysis team began by identifying a wide range of options but determined that some were technically or legally infeasible.  They then evaluated the remaining ones in more detail.
This included creating quantitative models to determine how decommissioning would affect fish production and ocean access.  They constructed a computer program that lets a user update the scores and weights.  They conducted sensitivity analysis to determine how uncertainties in costs and the impact of changing the weights on the attributes.  The most influential factor was the weight on compliance; a higher weight on compliance increased the desirability of complete removal.

After completing its analysis, the team then issued its report to its client and released its model to the stakeholders.  The deliberative part of the process included a series of meetings with stakeholders and the public and policy discussions that led to legislation that enables partial removal, the alternative that reduces both environmental impacts and costs.

I recommend the article as a case study of the analytic-deliberative process and an illustration of how decision analysis tools can be used.