The demand for numerical data on human rights (and human wrongs) has never been higher. Today human rights workers (i.e., aid workers, human rights activists, policy-makers and other practitioners) use numbers to show patterns of abuses, argue for funding, set priorities and figure out what works in service provision and advocacy. The United Nations Security Council has specifically mandated the collection of data on sexual violence during wartime. Advocacy organizations have increasingly deployed numerical data to make the case for attention and funding to particular emergencies. However, while data-driven approaches to human rights are laudable in many contexts, not all data are created equal. Some data illuminate; others just mislead.
The biggest problem with violence data is its uncertain relationship to true patterns of violence. For example, the reported murder rate declined pretty dramatically in my West Philadelphia neighborhood over the last ten years. This might mean that the true murder rate has fallen. Or, it might mean that people have stopped reporting crimes to the police, or that authorities are cooking the books. Without detailed knowledge about how crime reporting in my neighborhood has changed, it’s impossible to say which of these scenarios is accurate (or whether it’s a mix of all three, or something else entirely). But we often use numerical data in situations where we don’t have detailed knowledge of the context—what can be done?
Human rights workers can start by examining what kind of data they have access to. There are, roughly speaking, three kinds of data: statistical inferences, expert guesses and lists. Statistical inferences come from either systematically sampled survey data or multiple systems estimation. With a well-designed survey, the relationship between the population (for example, all households in my neighborhood) and the sample (for example, 100 randomly selected households in my neighborhood in 2006; another 100 randomly selected households in 2016) is known. If we assume that households are equally available to answer a survey, and are likely to give correct answers about whether someone in the household was murdered, then we can make a rigorous statistical inference (or estimate) of the true murder rate the neighborhood. Of course, these are big assumptions, ones that often aren’t met in violent contexts.