Showing posts with label bigdata. Show all posts
Showing posts with label bigdata. Show all posts

Monday, July 22, 2019

Algorithms and Auditability

@ruchowdh objects to an article by @etzioni and @tianhuil calling for algorithms to audit algorithms. The original article makes the following points.
  • Automated auditing, at a massive scale, can systematically probe AI systems and uncover biases or other undesirable behavior patterns. 
  • High-fidelity explanations of most AI decisions are not currently possible. The challenges of explainable AI are formidable.  
  • Auditing is complementary to explanations. In fact, auditing can help to investigate and validate (or invalidate) AI explanations.
  • Auditable AI is not a panacea. But auditable AI can increase transparency and combat bias. 

Rumman Chowdhury points out some of the potential imperfections of a system that relied on automated auditing, and does not like the idea that automated auditing might be an acceptable substitute for other forms of governance. Such a suggestion is not made explicitly in the article, and I haven't seen any evidence that this was the authors' intention. However, there is always a risk that people might latch onto a technical fix without understanding its limitations, and this risk is perhaps what underlies her critique.

In a recent paper, she calls for systems to be "taught to ignore data about race, gender, sexual orientation, and other characteristics that aren’t relevant to the decisions at hand". But how can people verify that systems are not only ignoring these data, but also being cautious about other data that may serve as proxies for race and class, as discussed by Cathy O'Neil? How can they prove that a system is systematically unfair without having some classification data of their own?

And yes, we know that all classification is problematic. But that doesn't mean being squeamish about classification, it just means being self-consciously critical about the tools you are using. Any given tool provides a particular lens or perspective, and it is important to remember that no tool can ever give you the whole picture. Donna Haraway calls this partial perspective.

With any tool, we need to be concerned about how the tool is used, by whom, and for whom. Chowdhury expects people to assume the tool will be in some sense "neutral", creating a "veneer of objectivity"; and she sees the tool as a way of centralizing power. Clearly there are some questions about the role of various stakeholders in promoting algorithmic fairness - the article mentions regulators as well as the ACLU - and there are some major concerns that the authors don't address in the article.

Chowdhury's final criticism is that the article "fails to acknowledge historical inequities, institutional injustice, and socially ingrained harm". If we see algorithmic bias as merely a technical problem, then this leads us to evaluate the technical merits of auditable AI, and acknowledge its potential use despite its clear limitations. And if we see algorithmic bias as an ethical problem, then we can look for various ways to "solve" and "eliminate" bias. @juliapowles calls this a "captivating diversion". But clearly that's not the whole story.

Some stakeholders (including the ACLU) may be concerned about historical and social injustice. Others (including the tech firms) are primarily interested in making the algorithms more accurate and powerful. So obviously it matters who controls the auditing tools. (Whom shall the tools serve?)

What algorithms and audits have in common is that they deliver opinions. A second opinion (possibly based on the auditing algorithm) may sometimes be useful - but only if it is reasonably independent of the first opinion, and doesn't entirely share the same assumptions or perspective. There are codes of ethics for human auditors, so we may want to ask whether automated auditing would be subject to some ethical code.




Paul R. Daugherty, H. James Wilson, and Rumman Chowdhury, Using Artificial Intelligence to Promote Diversity (Sloan Management Review, Winter 2019)

Oren Etzioni and Michael Li, High-Stakes AI Decisions Need to Be Automatically Audited (Wired, 18 July 2019)

Donna Haraway, Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective. In Simians, Cyborgs and Women (Free Association, 1991)

Cathy O'Neil, Weapons of Math Destruction

Julia Powles, The Seductive Diversion of ‘Solving’ Bias in Artificial Intelligence (7 December 2018)

Related posts: Whom Does the Technology Serve? (May 2019), Algorithms and Governmentality (July 2019)

Saturday, July 13, 2019

Algorithms and Governmentality

In the corner of the Internet where I hang out, it is reasonably well understood that big data raises a number of ethical issues, including data ownership and privacy.

There are two contrasting ways of characterizing these issues. One way is to focus on the use of big data to target individuals with increasingly personalized content, such as precision nudging. Thus mass surveillance provides commercial and governmental organizations with large quantities of personal data, allowing them to make precise calculations concerning individuals, and use these calculations for the purposes of influence and control.

Alternatively, we can look at how big data can be used to control large sets or populations - what Foucault calls governmentality. If the prime job of the bureaucrat is to compile lists that could be shuffled and compared (Note 1), then this function is increasingly being taken over by the technologies of data and intelligence - notably algorithms and so-called big data.

Although Deleuze challenges this dichotomy.
We no longer find ourselves dealing with the mass/individual pair. Individuals have become 'dividuals' and masses, samples, data, markets, or 'banks'.

Foucault's version of Bentham's panopticon is often invoked in discussions of mass surveillance, but what was equally important for Foucault was what he called biopower - a type of power that presupposed a closely meshed grid of material coercions rather than the physical existence of a sovereign. [Foucault 2003 via Adams]

People used to talk metaphorically about faceless bureaucracy being a machine, but now we have a real machine, performing the same function with much greater efficiency and effectiveness. And of course, scale.
The machine tended increasingly to dictate the purpose to be served, and to exclude other more intimate human needs. Lewis Mumford

Bureaucracy is usually regarded as a Bad Thing, so it's worth remembering that it is a lot better than some of the alternatives. Bureaucracy should mean you are judged according to an agreed set of criteria, rather than whether someone likes your face or went to the same school as you. Bureaucracy may provide some protection against arbitrary action and certain forms of injustice. And the fact that bureaucracy has sometimes been used by evil regimes for evil purposes isn't sufficient grounds for rejecting all forms of bureaucracy everywhere.

What bureaucracy does do is codify and classify, and this has important implications for discrimination and diversity.

Sometimes discrimination is considered to be a good thing. For example, recruitment should discriminate between those who are qualified to do the job and those who are not, and this can be based either on a subjective judgement or an agreed standard. But even this can be controversial. For example, the College of Policing is implementing a policy that police recruits in England and Wales should be educated to degree level, despite strong objections from the Police Federation.

Other kinds of discrimination such as gender and race are widely disapproved of, and many organizations have an official policy disavowing such discrimination, or affirming a belief in diversity. Despite such policies, however, some unofficial or inadvertent discrimination may often occur, and this can only be discovered and remedied by some form of codification and classification. Thus if campaigners want to show that firms are systematically paying women less than men, they need payroll data classified by gender to prove the point.

Organizations often have a diversity survey as part of their recruitment procedure, so that they can monitor the numbers of recruits by gender, race, religion, sexuality, disability or whatever, thereby detecting any hidden and unintended bias, but of course this depends on people's willingness to place themselves in one of the defined categories. (If everyone ticks the prefer not to say box, then the diversity statistics are not going to be very helpful.)

Daugherty, Wilson and Chowdhury call for systems to be taught to ignore data about race, gender, sexual orientation, and other characteristics that aren’t relevant to the decisions at hand. But there are often other data (such as postcode/zipcode) that are correlated with the attributes you are not supposed to use, and these may serve as accidental proxies, reintroducing discrimination by the back door. The decision-making algorithm may be designed to ignore certain data, based on training data that has been carefully constructed to eliminate certain forms of bias, but perhaps you then need a separate governance algorithm to check for any other correlations.

Bureaucracy produces lists, and of course the lists can either be wrong or used wrongly. For example, King's College London recently apologized for denying access to selected students during a royal visit.

Big data also codifies and classifies, although much of this is done on inferred categories rather than declared ones. For example, some social media platforms infer gender from someone's speech acts (or what Judith Butler would call performativity). And political views can apparently be inferred from food choice. The fact that these inferences may be inaccurate doesn't stop them being used for targetting purposes, or population control.

Cathy O'Neil's statement that algorithms are opinions embedded in code is widely quoted. This may lead people to think that this is only a problem if you disagree with these opinions, and that the main problem with big data and algorithmic intelligence is a lack of perfection. For example, criticizing such technologies as affective computing (to detect emotional state) if they fail to deal with ethnic diversity.

And of course technology companies encourage ethics professors to look at their products from this perspective, firstly because they welcome any ideas that would help them make their products more powerful, and secondly because it distracts the professors from the more fundamental question as to whether they should be doing things like facial recognition in the first place. @juliapowles calls this a "captivating diversion".

But a more fundamental question concerns the ethics of codification and classification. Following a detailed investigation of this topic, published under the title Sorting Things Out, Bowker and Star conclude that "all information systems are necessarily suffused with ethical and political values, modulated by local administrative procedures" (p321).
Black boxes are necessary, and not necessarily evil. The moral questions arise when the categories of the powerful become the taken for granted; when policy decisions are layered into inaccessible technological structures; when one group's visibility comes at the expense of another's suffering. (p320)
At the end of their book (pp324-5), they identify three things they want designers and users of information systems to do. (Clearly these things apply just as much to algorithms and big data as to older forms of information system.)
  • Firstly, allow for ambiguity and plurality, allowing for multiple definitions across different domains. They call this recognizing the balancing act of classifying.
  • Secondly, the sources of the classifications should remain transparent. If the categories are based on some professional opinion, these should be traceable to the profession or discourse or other authority that produced them. They call this rendering voice retrievable.
  • And thirdly, awareness of the unclassified or unclassifiable other. They call this being sensitive to exclusions, and note that residual categories have their own texture that operates like the silences in a symphony to pattern the visible categories and their boundaries (p325).




Note 1: This view is attributed to Bruno Latour by Bowker and Star (1999 p 137). However, although Latour talks about paper-shuffling bureaucrats (1987 pp 254-5), I have been unable to find this particular quote.

Rachel Adams, Michel Foucault: Biopolitics and Biopower (Critical Legal Thinking, 10 May 2017)

Geoffrey Bowker and Susan Leigh Star, Sorting Things Out (MIT Press 1999).

Paul R. Daugherty, H. James Wilson, and Rumman Chowdhury, Using Artificial Intelligence to Promote Diversity (Sloan Management Review, Winter 2019)

Gilles Deleuze, Postscript on the Societies of Control (October, Vol 59, Winter 1992), pp. 3-7

Michel Foucault, ‘Society Must be Defended’ Lecture Series at the Collège de France, 1975-76 (2003) (trans. D Macey)

Maša Galič, Tjerk Timan and Bert-Jaap Koops, Bentham, Deleuze and Beyond: An Overview of Surveillance Theories from the Panopticon to Participation (Philos. Technol. 30:9–37, 2017)

Bruno Latour, Science in Action (Harvard University Press 1987)

Lewis Mumford, The Myth of the Machine (1967)

Samantha Murphy, Political Ideology Linked to Food Choices (LiveScience, 24 May 2011)

Julia Powles, The Seductive Diversion of ‘Solving’ Bias in Artificial Intelligence (7 December 2018)

Antoinette Rouvroy and Thomas Berns (translated by Elizabeth Libbrecht), Algorithmic governmentality and prospects of emancipation (Réseaux No 177, 2013)

BBC News, All officers 'should have degrees', says College of Policing (13 November 2015), King's College London sorry over royal visit student bans (4 July 2019)


Related posts

Quotes on Bureaucracy (June 2003), Crude Categories (August 2009), What is the Purpose of Diversity? (January 2010), Affective Computing (March 2019), The Game of Wits between Technologists and Ethics Professors (June 2019), Algorithms and Auditability (July 2019), On the Performativity of Data (August 2021), The Corporate Sorting Hat (September 2021)


Updated 16 July 2019

Wednesday, January 30, 2013

Real Criticism, The Subject Supposed to Know

"Goodbye, Anecdotes", says @Butterworthy, "The Age Of Big Data Demands Real Criticism" (AWL, January 2013). Thanks to @milouness, who comments "Important concepts here about what is knowable!".  The article tries to link Big Data with Big Questions about the Big Picture, and what @Butterworthy calls The Big Criticism. From this perspective, Bill Franks' advice, To Succeed with Big Data, Start Small (HBR Oct 2012), is downright paradoxical.

But why would we expect Big Data to help us answer the Big Questions? Big Data is rather a misnomer: it mostly comprises very large quantities of very small data and very weak signals. Retailers wade through Big Data in order to fine-tune their pricing strategies; pharma researchers wade through Big Data in order to find chemicals with a marginal advantage over some other chemicals; intelligence analysts wade through Big Data to detect terrorist plots. Doubtless these are useful and sometimes profitable exercises, but they are hardly giving us much of a Big Picture. Big Data may give us important clues about what the terrorists are up to, but it doesn't tell us why.

A few years ago, Chris Anderson promoted The End of Theory, and published an article claiming that The Data Deluge Makes the Scientific Method Obsolete (Wired June 2008), although this may have only been an ironic reference to Fukuyama's earlier idea of The End of History. Claiming obsolescence seems like hyperbole, although scientific method has always been modified by technological progress. Even in mathematics, computer power and human brilliance have combined to crack some previously unsolved problems. See for example, Proof and Beauty (The Economist, March 2005).

Although @Butterworthy claims to have identified some critical ("Big Critical") questions, there seems to be only one real question - the dialectical question of quantity becoming quality. Are we on the cusp of aggregating utilitarianism into new tyrannies of scale? Are the numbers are so big, they leave interpretation behind and acquire their own agency? How much information and of what kind would you need to conclude something - for example, something like gender bias in the media?

A recent academic study looked at 2.4 million pages of newspaper and came to the conclusion that there was some gender bias. That's a lot of newspaper. It's like examining every single grain of sand in the forest for traces of ursine faeces. (In other words, looking for microscopic proof that bears defecate in the woods.) From a technophile perspective, Big Data seems to be raising the bar for scientific methodology: following this impressive piece of research, those who don't understand the concept of statistical significance can dismiss any smaller study - for example, one that merely studied thousands of pages - as unscientific anecdote. At a stroke, decades of careful analysis by feminists can be discredited because their sample sizes were too small by modern Big Data standards, and so there is now less scientifically credible evidence of gender bias than there was before.

Seriously, how many pages of newspaper do you have to read to convince yourself of gender bias? Clearly this is an example of Big Data getting in the way of the Big Picture. @Butterworthy clearly understands this danger, and sees the redemptive possibility of Big Crit (whatever that is) revitalizing the notion of critical authority and restoring some balance to the universe. I'm not sure I follow how he thinks that is going to happen. 


 Related post: Big Data and Organizational Intelligence (November 2018)