AssociationAI / AI Literacy
Trihelix AI team Published

Article

Beginner

When AI Is the Wrong Tool: A Sorting Sheet Built From Four Studies

Four studies point to the same trouble spots: choices, member-facing answers, practice that builds skill and research tools sold as reliable. None studied an association.

AI Limits Task Selection Risk Research Board

A membership director is asked to approve AI for four jobs: answering member questions, ranking award nominees, training new staff, and researching a policy question for the board. The research suggests the answer should differ by job, and it comes from outside associations, because none of the four sources we read studied one.

The evidence does not say to stop using AI. One study below found greater gains for content creation, and another found its harm was largely mitigated once teachers designed the tool’s limits. What it supports is sorting jobs before choosing a tool, and the sheet below is our own way of doing that, not something any of the four sources proposes.

Diagram: four studies and who each one covered, which were people in experiments, one airline passenger, students at one high school and legal research queries, with association staff shown as not studied in any of them.

A sorting sheet for four kinds of job

Give every current or planned use of AI its own row, and answer the middle column before anything else. An honest “not sure” is treated as a yes.

Kind of jobQuestion to answer firstIf the answer is yes
A decision staff make (approve, rank, accept, refuse)Does the job end in a choice between outcomes?Run past cases through staff alone, the tool alone and both together, and adopt the pairing only if it beats the better of the first two.
An answer a member will rely onCould a member pay, register, cancel or miss a deadline because of it?Publish only text a staff member owns, and point the member to a page staff maintain.
Practice that builds a skillWill this person have to do the job without the tool later?Have the person attempt it first, and give the tool a role as a hint-giver, not an answer-giver.
Research the board will act onIs the tool sold as accurate or free of made-up answers?Ask for the test behind the claim, and check every cited fact by hand before it reaches the board.

Row one: choices are where the pairing slipped

The abstract of a Nature Human Behaviour meta-analysis by MIT researchers reports a preregistered review of 106 experiments, published from January 2020 to June 2023, in which participants did tasks alone, with AI, or left them to AI. On average, the same abstract says, the pairing scored below the better of the two working alone, judgment calls were where it slipped, and drafting and other content creation was where it showed greater gains. It also reports that the pairing gained where people alone beat the AI and lost where the AI alone beat people, and the authors name possible publication bias and uneven study designs as limits.

The participants were not association staff, and the studies were published by June 2023, before the tools staff use now, so we cannot say whether the average has moved. We read the result as a reason to measure rather than assume. That is why the first row asks for a test on your own past cases.

Row two: what the chatbot said became the organization’s statement

A decision of the British Columbia Civil Resolution Tribunal, a provincial government body ruling on a small claim from one airline passenger, found that the airline’s website chatbot had suggested the passenger could claim a bereavement fare after travelling. In its reasons, the tribunal treated the chatbot as part of the airline’s website, said the airline carried the same responsibility as for any static page, found it had not taken reasonable care to make the chatbot accurate, and awarded the passenger $650.88 in damages.

This is one small-claims ruling from British Columbia, and nothing we read shows how a United States court would treat the same facts. The lesson for the sheet is narrower: an answer that a member can act on is a statement from your organization, and someone has to own it. The how-to side lives in a fact-check routine for member-facing AI and a guide to member FAQ bots that keep data private.

Row three: practice with the answers on hand

A Proceedings of the National Academy of Sciences paper by University of Pennsylvania researchers and collaborators reports a randomized controlled trial in a Turkish high school, covering about fifty mathematics classes and nearly 1,000 students. In that trial, students who practiced with a standard chat interface to a commercial language model did better on the practice problems but worse than classmates who never had it once it was taken away, and a version that gave teacher-designed hints instead of answers largely removed the harm.

The authors point out that the results come from one subject at one school. Teenagers learning mathematics are not adults learning a membership system, so we treat the study as a warning about how the tool is set up and not as a finding about staff. If a new hire is meant to learn how to read a dues report, the sheet says to let them try first.

Row four: tools described as reliable

The abstract of a Journal of Empirical Legal Studies article by Stanford and Yale researchers reports a preregistered test of commercial legal research tools that their providers had described as eliminating or avoiding made-up answers. That abstract calls the providers’ claims overstated and says each tool tested made up an answer roughly one time in six to one time in three on legal queries, though less often than general-purpose chatbots.

Legal research is not association research, and the tools may have changed since the test, which the abstract cannot tell us. Still, a tool being sold as reliable is not the same as a tool being shown to be reliable, so the fourth row asks for the test behind the claim. A guide to evaluating a proprietary platform’s claims covers what to ask.

Where the sheet can mislead

The sheet covers four kinds of job because those are the four the evidence touches, and it leaves out plenty of work where nobody has measured anything. A “yes” also does not mean “never use AI”. Every row still allows a drafted note, a hint or a first search, and the meta-analysis found greater gains in content creation. What a “yes” asks for is a person with the final say and a check that someone has actually run.

The school study shows how much setup matters: the same trial found one configuration harmed learning while another mostly did not, so a tool’s settings can matter as much as the tool. The other risk is false comfort from the “no” column. The studies cover experiments, a school, a tribunal and legal queries, and a job that looks safe on the sheet can still go wrong in ways no study here describes. For jobs marked no, a short trial that runs the same input three times shows how much the output moves before anyone relies on it.

Fill the sheet in once with the people who do the work, then ask the director to sign off on the rows marked yes. The first question to settle at that meeting is who owns each of those rows, and the second is what check will be run.

Sources (4)