Monideepa Tarafdar, PhD, Isenberg’s Charles J. Dockendorff endowed professor, recently published research in the Journal of Management Information Systems focused on answering the question that a
Monideepa Tarafdar

Monideepa Tarafdar, PhD, Isenberg’s Charles J. Dockendorff endowed professor, recently published research in the Journal of Management Information Systems focused on answering the question that all business leaders are asking in the age of artificial intelligence: How can we work effectively with machines that appear to think—such as, in today’s world, large language models?

“Research shows that humans can over-rely on such machines or experience aversion to them,” Tarafdar said, explaining that both reactions to large language models (LLMs) are likely to lead to bad outcomes. People who accept answers from LLMs without critical evaluation end up with confident and polished-sounding material that might be “wrong, biased, hallucinated, or irrelevant,” she said. Those who don’t trust LLMs to be helpful, on the other hand, miss out on support that could help improve their work. 

Tarafdar and her co-authors wanted to investigate how humans can increase the quality of LLM outputs, so they  focused on the concept of reflection, which she defined as “taking a hard look at what you already know about a topic, then comparing it with what the LLM provides,” she said. “It’s about weighing the two side by side, spotting where they align or clash, and using that comparison to build a deeper, more refined understanding of the topic for yourself.” 

Tarafdar answered some questions about the research:

How did you study reflection in this context? 

We explore two ideas relating to reflection. 

The internal aspect—how people think while using LLMs—is captured in three modes of thinking: shallow, adaptive, and dialogic. We created ways to measure these modes so we could see how they play out in practice. The shallow mode is when the human relies only on their existing knowledge to tackle the task, without updating it. The adaptive mode is when the individual tests their current understanding against new information, making small adjustments without fully questioning their core assumptions. It allows for flexibility and incremental improvement while still staying within familiar ways of thinking. The dialogic is when the human critically challenges their own assumptions, using new information to upend their understanding. It often leads to transformative learning, where old perspectives are overturned.

The external aspect—how people interact with LLMs—is captured in two interaction styles: conversational versus adversarial. The first is where the AI’s answers line up with what we already believe, and second is where they challenge or contradict us.

We explore how the combination of these three states and two interacting styles affects task factuality.

By combining these two angles—mode of thinking and interaction style—we were able to take a detailed, data‑driven look at reflection. This approach helps to tease apart what comes from the user (their thinking mode) and what can be shaped by design (prompts that encourage a certain interaction style).

Finally, we placed reflection into a bigger framework that connects it directly to task accuracy. In other words, we showed how the way we think and the way we interact with AI can predict how good the task output will be.

Your study shows that taking certain approaches in working with an LLM to complete a task leads to more factually accurate results—can you break down what the best approach is?

It turns out that how we think (our internal thinking mode) and how the LLM interacts with us (interaction mode) both play a role in whether the output is accurate and trustworthy. We found that not all thinking modes bear the same fruit. The real game‑changer happens when we combine two things:

  • Adversarial interaction: Instead of just accepting what we say, the LLM challenges us and pushes back on what we say.
  • Dialogic thinking: This is the thinking mode with the deepest level of engagement, where we stay alert to what we know, and yet are open to reshaping our understanding based on the LLM’s inputs.

Together, these two create the strongest boost in factual accuracy. In simple terms: the best results come when you stay sharp on the inside and the AI challenges you on the outside.

 

In your experience as a researcher and as an instructor, what are the most common mistakes people are making in their interactions with LLMs?

There are two issues. One is a lack of cognitive vigilance while using LLMs, which means that people often don’t dig very deep into the information they get. Instead, they tend to skim it, accept it at face value, and move on. This kind of shallow engagement can lead to weak arguments, missing details, missed opportunities for learning, and mistakes in the final work. 

The second and related one is lack of cognitive primacy, which means that the human does not take the lead in their use of LLMs. Studies show that people struggle to judge whether an LLM’s output is reliable—sometimes they trust it too much, and other times they dismiss it too quickly. In both cases, they are not using LLMs in a way that leverages the capacities of both the human and the LLM. Through the concept of reflection we have tried to walk this rather thin line. We are essentially saying, be open to the LLM’s outputs but also foreground what you yourself know.

 

Did you come away from this research with ideas about how AI companies can encourage these more productive approaches for human users?

Unlike most current LLM designs, which tend to simply follow the user’s lead, and even lead to sycophantic interactions, we suggest building LLMs that push back at the human. Instead of just agreeing, they should encourage users to question, challenge, and dig deeper into the outputs. This can be done at the code level and also by explicitly prompting the user to ask the LLM to be a ‘sparring partner’ rather than a ‘friend.’ 

Essentially, LLM users, when they are processing LLMs’ outputs to their questions and brainstorming with LLMs, should tell the LLMs, “Say as I don’t say.” This way, they will stay actively engaged, will be prepared for surprising answers, and are more likely to spot errors or gaps in their own understanding as well as the LLM’s output. Companies should train employees to function in this mode. 

Here’s some practical advice: If you want more reliable results from AI, don’t just nod along. Stay curious, stay critical, and ask the machine to argue with you.

 

What are some contexts in which you personally use LLMs?

I use LLMs primarily for searching literature and information, research methods related support and brainstorming  (e.g., curriculum and assessment preparation). I also incorporate LLMs into course assignments where the goal is to learn how to use them responsibly for everyday managerial tasks such as creating promotion campaigns, running data analysis and creating reports. 

 

What steps do you take to ensure that your own results are factually accurate?

One of the learnings from this research has been that I actively and mindfully try to consider LLM outputs with an open mind while relying on my own knowledge and judgment to help me identify gaps in my own understanding and also push back against any inconsistencies or weaknesses I perceive in LLMs’ outputs.

 

Do you have an opinion about whose responsibility it is (or should be) to ensure that end results of human-LLM collaborations are factually accurate?

The responsibility rests on all stakeholders. LLM (tech) companies should design LLMs that encourage users to dig deeper, LLM users should challenge the outputs they receive from the LLM as well as learn from them and organizations should build a culture where employees are motivated to remain open to LLMs’ outputs as well as be confident about their own understanding in their interactions with LLMs.