Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Sunday, July 6, 2025

How Soon Might Humans Be Replaced At Work

As noted by Thomas Claburn in The Register, there seems to be a contradiction between two pieces of research relating to the development and use of AI in business organizations.

On the one hand, teams of researchers have developed benchmarks to study the effectiveness of AI, and have found success rates between 25% and 40%, depending on the situation.

On the other hand, Gartner reports that business executives are expecting a success rate nearer to 60% - if we interpret not-being-cancelled as a marker for success. More than 40 percent of agentic AI projects will be cancelled by the end of 2027 due to rising costs, unclear business value, or insufficient risk controls.

 

History tells us that the adoption of technology to perform work is only partially dependent on the quality of the work, and can often be driven more by cost. The original Luddites protested at the adoption of machines to replace textile workers, but their argument was largely based on the inferior quality of the textiles produced by the machines. It was only later that this label was attached to anyone who resisted technology on principle.

Around ten years ago, I attended a debate on artificial intelligence sponsored by the Chartered Institute of Patent Agents. In my commentary on this debate (How Soon Might Humans Be Replaced At Work?) I noted that decision-makers may easily be tempted by short-term cost savings from automation, even if the poor quality of the work results in higher costs and risks in the longer term.

In their look at the labour market potential of AI, Tyna Eloundou et al note that

A key determinant of their utility is the level of confidence humans place in them and how humans adapt their habits. For instance, in the legal profession, the models’ usefulness depends on whether legal professionals can trust model outputs without verifying original documents or conducting independent research. ... Consequently, a comprehensive understanding of the adoption and use of LLMs by workers and firms requires a more in-depth exploration of these intricacies.

However, while levels of confidence and trust can be assessed by surveying people's opinions, such surveys cannot assess whether these levels of confidence and trust are justified. Graham Neubig told The Register that this was what prompted the development of a more objective benchmark for AI effectiveness.


Thomas Claburn, AI agents get office tasks wrong around 70% of the time, and a lot of them aren't AI at all (The Register, 29 June 2025)

Thomas Claburn, AI has had zero effect on jobs so far, says Yale study (The Register, 1 October 2025)

Tyna Eloundou, Sam Manning, Pamela Mishkin and Daniel Rock, GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models (August 2023)

Wikipedia: Luddite 

Related Posts: How Soon Might Humans Be Replaced At Work? November 2015, RPA - Real Value or Painful Experimentation? (August 2019), Explaining Layoffs (October 2025)

Sunday, November 3, 2024

Influencing the Habermas Machine

In my previous post Towards the Habermas Machine, I talked about a large language model (LLM) developed by Google DeepMind for generating a consensus position from a collection of individual views, named after Jürgen Habermas.

Given that democratic deliberation relies on knowledge of various kinds, followers of Habermas might be interested in how knowledge is injected into discourse. Habermas argued that mutual understanding was dependent upon a background stock of cultural knowledge that is always already familiar to agents. but this clearly has to be supplemented by knowledge about the matter in question.

For example, we might expect a discussion about appropriate speed limits to be informed by reliable or unreliable beliefs about the effects of a given speed limit on journey times, accident rates, pollution, and so on. In traditional discussion forums, it is extremely common for people to present themselves as having some special knowledge or authority, which supposedly gives extra weight to their opinions, and we might expect something similar to happen in a tech-enabled version.

For many years, the Internet has been distorted by Search Engine Optimization (SEO), which means that the results of an internet search are largely driven by commercial interests of various kinds. Researchers have recently raised a similar issue in relation to large language models, namely Generative Engine Optimization (GEO). Meanwhile, other researchers have found that LLMs (like many humans) are more impressed by superficial jargon than by proper research.

So we might reasonably assume that various commercial interests (car manufacturers, insurers, oil companies, etc) will be looking for ways to influence the outputs of the Habermas Machine on the speed limit question by overloading the Internet with knowledge (regime of truth) in the appropriate format. Meanwhile the background stock of cultural knowledge is now presumably co-extensive with the entire Internet.

Is there anything that the Habermas Machine can do to manage the quality of the knowledge used in its deliberations?


Footnote: Followers of Habermas can't agree on the encyclopedia entry, so there are two rival versions.

Footnote: The relationship between knowledge and discourse goes much wider than Habermas, so interest in this question is certainly not limited to his followers. I might need to write a separate post about the Foucault Machine.


Pranjal Aggarwal et al, GEO: Generative Engine Optimization (arxiv v3, 28 June 2024)

Callum Bains, The chatbot optimisation game: can we trust AI web searches? (Observer, 3 November 2024)

Alexander Wan, Eric Wallace, Dan Klein, What Evidence Do Language Models Find Convincing? (arxiv v2, 9 August 2024)

Stanford Encyclopedia of Philosophy: Jürgen Habermas (v1 2007) Jürgen Habermas (v2 2023)

Saturday, October 19, 2024

Towards the Habermas Machine

Google DeepMind has just announced a large language model, which claims to generate a consensus position from a collection of individual views. The name of the model is a reference to Jürgen Habermas’s theory of communicative action.

An internet search for Habermas machine throws up two previous initiatives under the same name. Firstly an art project by Kristopher Holland.

The Habermas Machine (2006–2012) is a conceptual art experience that both examines and promotes an experiential relation to Jürgen Habermas’ grand theory for understanding human interaction. The central claim is that The Theory of Communicative Action can be experienced, reflected upon and practised when encountered within arts-based research. Habermas’ description of how our everyday lives are founded by intersubjective experience, and caught up in certain normative, objective and subjective contexts is transformed through the method of conceptual art into a process of collaborative designing, enacting and articulating. This artistic reframing makes it possible to experience the communicative structure of knowledge and the ontological structure of intersubjectivity in a practice of non-discursive ‘philosophy without text’. Feiten Holland Chemero

And secondly, an approach to Dialogue Mapping described as a device that all participants can climb into and converse with complete communicative rationality, contained in a book by @paulculmsee and Jailash Awati, and mentioned in this Reddit post Why is Dialogue Mapping not wide spread? Dialogue Mapping was developed by Jeff Conklin and others as an approach to addressing wicked problems. See also Issue Based Information Systems (IBIS).


Update

Christopher Summerfield, one of the authors of the DeepMind paper, spoke at the Royal Society on October 29th 2024. https://www.youtube.com/live/cW1Wq7_8v1Y?si=oqo8Lw7479x4QqKt&t=18890

All the examples shown in his talk were policy matters that could be reduced to Yes/No questions. Such questions would traditionally be surveyed by asking people to place themselves on a scale from Strongly Agree to Strongly Disagree, and it is easy to see how a language-based method such as the Habermas Machine offers some advantages over a numerical scale. But not clear how this works for more provocative questions, let alone wicked problems.

Someone in the audience asked if this method would work in what he called a compromised democracy, and Summerfield acknowledged that the method assumes what he called a good faith scaffold. Obviously all democracies in the real world are imperfect, and he didn't go into the question as to how sensitive or vulnerable the method might be to such imperfections, but the method might conceivably help to overcome some of these imperfections under some conditions: for example, Summerfield referred specifically to the tyranny of the majority.

While the performance of the Habermas machine in their study compared favourably with the performance of human mediators, Summerfield suggested that we should move away from thinking about AI in these terms. The point is not to create AI-based agents that can behave like intelligent people but to build intelligent institutions - tools for creating social order and fostering cooperation. As my regular readers will know, orgintelligence has long been an important theme for this blog. See for example my post On Organizations and Machines (January 2022).


Jeffrey Conklin, Dialogue Mapping: Building Shared Understanding of Wicked Problems (Wiley 2006). See also CogNexus website.

Paul Culmsee and Jailash Awati, The Heretic's Guide to Best Practice (2013)

Nicola Davis, AI mediation tool may help reduce culture war rifts, say researchers (Guardian, 17 October 2024)

Tim Elmo Feiten, Kristopher Holland and Anthony Chemero, Doing philosophy with a water-lance: art and the future of embodied cognition (Adaptive Behavior 2021) 

Michael Tessler et al, AI can help humans find common ground in democratic deliberation (Science, 18 October 2024)

Beyond the symbols vs signals debate (The Royal Society, 28-29 October 2024)

Wikipedia: Issue Based Information Systems (IBIS), Wicked Problem

See also Influencing the Habermas Machine (November 2024)

Tuesday, September 26, 2023

Creativity and Recursivity

Prompted by @jjn1's article on AI and creative thinking, I've been reading a paper by some researchers comparing the creativity of ChatGPT against their students (at an elite university, no less).

What is interesting about this paper is not that ChatGPT is capable of producing large quantities of ideas much more quickly than human students, but that the evaluation method used by the researchers rated the AI-generated ideas as being of higher quality. From 200 human-generated ideas and 200 algorithm-generated ideas, 35 of the top-scoring 40 were algo-generated.

So what was this evaluation method? They used a standard market research survey, conducted with college-age individuals in the United States, mediated via mTurk. Two dimensions of quality were considered: purchase intent (would you be likely to buy one) and novelty. The paper explains the difficulty of evaluating economic value directly, and argues that purchase intent provides a reasonable indicator of relative value.

The paper discusses the production cost of ideas, but this doesn't tell us anything about what the ideas might be worth. If ideas were really a dime a dozen, as the paper title suggests, then neither the impressive productivity of ChatGPT nor the effort of the design students would be economically justified. But the production of the initial idea is only a tiny fraction of the overall creative process, and (with the exception of speculative bubbles) raw ideas have very little market value (hence dime a dozen). So this research is not telling us much about creativity as a whole.

A footnote to the paper considers and dismisses the concern that some of these mTurk responses might have been generated by an algorithm rather than a human. But does that algo/human distinction even hold up these days? Most of us nowadays inhabit a socio-technical world that is co-created by people and algorithms, and perhaps this is particularly true of the Venn diagram intersection between college-age individuals in the United States and mTurk users. If humans and algorithms increasingly have access to the same information, and are increasingly judging things in similar ways, it is perhaps not surprising that their evaluations converge. And we should not be too surprised if it turns out that algorithms have some advantages over humans in achieving high scores in this constructed simulation.

(Note: Atari et al recommend caution in interpreting comparisons between humans and algorithms, as they argue that those from Western, Educated, Industrialized, Rich and Democratic societies - which they call WEIRD - are not representative of humanity as a whole.)

A number of writers on algorithms have explored the entanglement between humans and technical systems, often invoking the concept of recursivity. This concept has been variously defined in terms of co-production (Hayles), second-order cybernetics and autopoiesis (Clarke), and being outside of itself (ekstasis), which recursively extends to the indefinite (Yuk Hui). Louise Amoore argues that, in every singular action of an apparently autonomous system, then, resides a multiplicity of human and algorithmic judgements, assumptions, thresholds, and probabilities.

(Note: I haven't read Yuk Hui's book yet, so his quote is taken from a 2021 paper)

Of course, the entanglement doesn't only include the participants in the market research survey, but also students and teachers of product design, yes even those at an elite university. This is not to say that any of these human subjects were directly influenced by ChatGPT itself, since much of the content under investigation predated this particular system. What is relevant here is algorithmic culture in general, which as Ted Striphas's new book makes clear has long historical roots. (Or should I say rhizome?)

What does algorithmic culture entail for product design practice? For one thing, if a new product is to appeal to a market of potential consumers, it generally has to achieve this via digital media - recommended by algorithms and liked by people (and bots) on social media. Thus successful products have to submit to the discipline of digital platforms: being sorted, classified and prioritized by a complex sociotechnical ecosystem. So we might expect some anticipation of this (conscious or otherwise) to be built into the design heuristics (or what Peter Rowe, following Gadamer, calls enabling prejudices) taught in the product design programme at an elite university.

So we need to be careful not to interpret this research finding as indicating a successful invasion of the algorithm into a previously entirely human activity. Instead, it merely represents a further recalibration of algorithmic culture in relation to an existing sociotechnical ecosystem. 


Update April 2024

As far as I can see, the evaluation method used in this study did not consider the question of feasibility. If students have a stronger sense of the possible than algorithms do, this may inhibit their ability to put forward superficially attractive but practically ridiculous ideas, which might nevertheless score highly on the evaluation method used here. In my post ChatGPT and Entropy (April 2024), I look at the phenomenon of model collapse, which could lead to algorithms becoming increasingly disconnected from reality. But perhaps able to generate increasingly outlandish ideas?

 


Louise Amoore, Cloud Ethics: Algorithms and the Attributes of Ourselves and Others (Durham and London: Duke University Press 2020)

Mohammad Atari, Mona J. Xue, Peter S. Park, Damián E. Blasi and Joseph Henrich, Which Humans? (PsyArXiv, September 2023) HT @MCoeckelbergh

David Beer, The problem of researching a recursive society: Algorithms, data coils and the looping of the social (Big Data and Society, 2022)

Bruce Clarke, Rethinking Gaia: Stengers, Latour, Margulis (Theory Culture and Society 2017)

Karan Girotra, Lennart Meincke, Christian Terwiesch, and Karl T. Ulrich, Ideas are Dimes a Dozen: Large Language Models for Idea Generation in Innovation (10 July 2023)

N Katherine Hayles, The Illusion of Autonomy and the Fact of Recursivity: Virtual Ecologies, Entertainment, and "Infinite Jest" New Literary History , Summer, 1999, Vol. 30, No. 3, Ecocriticism (Summer, 1999), pp. 675-697

Yuk Hui, Problems of Temporality in the Digital Epoch, in Axel Volmar and Kyle Stine (eds) Media Infrastructures and the Politics of Digital Time (Amsterdam University Press 2021)

John Naughton, When it comes to creative thinking, it’s clear that AI systems mean business (Guardian, 23 September 2023) 

Peter Rowe, Design Thinking (MIT Press 1987)

Ted Striphas, Algorithmic culture before the internet (New York: Columbia University Press, 2023)

Richard Veryard, As We May Think Now (Subjectivity 30/4, 2023)

See also:  From Enabling Prejudices to Sedimented Principles (March 2013)

Saturday, February 18, 2023

Hedgehog Innovation

According to Archilochus, the fox knows many things, but a hedgehog knows one big thing.

In his article on AI and the threat to middle class jobs, Larry Elliot focuses on machine learning and robotics.

AI stands to be to the fourth industrial revolution what the spinning jenny and the steam engine were to the first in the 18th century: a transformative technology that will fundamentally reshape economies.

When people write about earlier waves of technological innovation, they often focus on one technology in particular - for example a cluster of innovations associated with the adoption of electrification in a wide range of industrial contexts.

While AI may be an important component of the fourth industrial revolution, it is usually framed as an enabler rather than the primary source of transformation. Furthermore, much of the Industry 4.0 agenda is directed at physical processes in agriculture, manufacturing and logistics, rather than clerical and knowledge work. It tends to be framed as many intersecting innovations rather than one big thing.

There is also a question about the pace of technological change. Elliott notes a large increase in the number of AI patents, but as I've noted previously I don't regard patent activity as a reliable indicator of innovation. The primary purpose of a patent is not to enable the inventor to exploit something, it is to prevent anyone else freely exploiting it. And Ezrachi and Stucke provide evidence of other ways in which tech companies stifle innovation.

However the AI Index Report does contain other measures of AI innovation that are more convincing.


 AI Index Report (Stanford University, March 2022)

Larry Elliott, The AI industrial revolution puts middle-class workers under threat this time (Guardian, 18 February 2023)

Ariel Ezrachi and Maurice Stucke, How Big-Tech Barons Smash Innovation and how to strike back (New York: Harper, 2022)

Wikipedia: Fourth Industrial Revolution, The Hedgehog and the Fox

Related Posts: Evolution or Revolution (May 2006), It's Not All About (July 2008), Hedgehog Politics (October 2008), The New Economics of Manufacturing (November 2015), What does a patent say? (February 2023)

Sunday, January 22, 2023

Reasoning with the majority - chatGPT

#ThinkingWithTheMajority 

#chatGPT has attracted considerable attention since its launch in November 2022, prompting concerns about the quality of its output as well as the potential consequences of widespread use and misuse of this and similar tools.

Virginia Dignum has discovered that it has a fundamental misunderstanding of basic propositional logic. In answer to her question, chatGPT claims that the statement "if the moon is made of cheese then the sun is made of milk" is false, and goes on to argue that "if the premise is false then any implication or conclusion drawn from that premise is also false". In her test, the algorithm persists in what she calls "wrong reasoning".

I can't exactly recall at what point in my education I was introduced to propositional calculus, but I suspect that most people are unfamiliar with it. If Professor Dignum were to ask a hundred people the same question, it is possible that the majority would agree with chatGPT.

In which case, chatGPT counts as what A.A. Milne once classified as a third-rate mind - "thinking with the majority". I have previously placed Google and other Internet services into this category.

Other researchers have tested chatGPT against known logical paradoxes. In one experiment (reported via LinkedIn) it recognizes the Liar Paradox when Epimenides is explicitly mentioned in the question, but apparently not otherwise. No doubt someone will be asking it about the baldness of the present King of France.

One of the concerns expressed about AI-generated text is that it might be used by students to generate coursework assignments. At the present state of the art, although AI-generated text may look plausible it typically lacks coherence and would be unlikely to be awarded a high grade, but it could easily be awarded a pass mark. In any case, I suspect many students produce their essays by following a similar process, grabbing random ideas from the Internet and assembling them into a semi-coherent narrative but not actually doing much real thinking.

There are two issues here for universities and business schools. Firstly whether the use of these services counts as academic dishonesty, similar to using an essay mill, and how this might be detected, given that standard plagiarism detection software won't help much. And secondly whether the possibility of passing a course without demonstrating correct and joined-up reasoning (aka "thinking") represents a systemic failure in the way students are taught and evaluated.


See also

Andrew Jack, AI chatbot’s MBA exam pass poses test for business schools (FT, 21 January 2023) HT @mireillemoret

Gary Marcus, AI's Jurassic Park Moment (CACM, 12 December 2022)

Christian Terwiesch, Would Chat GPT3 Get a Wharton MBA? (Wharton White Paper, 17 January 2023)

Related posts: Thinking with the Majority (March 2009), Thinking with the Majority - a New Twist (May 2021), Satanic Essay Mills (October 2021)

Wikipedia: ChatGPT, Entailment, Liar Paradox, Plagiarism, Propositional calculus