Fluent but Fragile
Preserving Professional Judgment in Classrooms and Operational Organizations
BLUF (Bottom Line Up Front): Generative AI can create polished writing before the user fully understands the source material. It may leave out important exceptions, reduce complex ideas, or hide contradictions while sounding more certain than the evidence supports. AI-assisted work should stay connected to sources that people can review and to a named owner who can explain and defend the final result.
Why I Am Writing This
I approach this issue as a college professor and senior project manager supporting enterprise and Department of Defense delivery. In both settings, I have seen polished drafts reach the review stage even though the listed author could not trace a claim to its source or explain what evidence might change the recommendation.
"Hallucination" is the most obvious problem. Large language models generate likely sequences of words, but they do not independently verify facts or resolve conflicts in the source material. As a result, a clear-sounding explanation may hide missing information or contradictions. In a long document, an important exception may also be buried deep in the record.
The immediate result is a loss of credibility because a polished first draft becomes irrelevant if the writer cannot answer a basic question about the evidence presented. Experienced employees face the same problem more gradually. AI-generated summaries reduce the time they spend working directly with the source material, while responsibility for the final product still remains with the human user.
A Classroom Policy for Intentional AI Use
In my courses, students need permission before using AI on an assignment. They let me know what they want to use it for and I provide them guardrails forcing the scope of use into the open before they begin their work. Search support may help a student locate material, but generating the argument bypasses the learning objective.
When a submission exceeds the student's demonstrated command of the subject, I ask for an oral defense. I may also request the prompt and first output. Tracked revisions indicate where the student's own reasoning was incorporated into the work and whether the weak points in the first output were corrected. Operational reviews can use the same standard; two questions can stop a weak product from advancing:
- Who verified it?
- Who owns the decision?
Probability Bias and the Limits of Model Understanding
AI summaries and drafts are attractive because they turn large source sets into readable prose. The human judgment involved in determining inclusions and conclusions is excluded while the content of the summary has been outsourced to the AI model to begin framing the recommendation before the reviewer sees the full record.
Confident language reduces the level of review, especially when an important exception is buried in a policy or engineering note, which can make the entire recommendation unusable.
Bender et al. (2021) explain that large language models create language by recognizing patterns in training data. They do not understand the subject through experience or make independent judgments about whether a statement is true. Ji et al. (2023) also show that AI-generated text can misrepresent its source material or add claims that the evidence does not support. This creates a serious risk in professional work. A draft may sound clear, current, and confident while still presenting the facts incorrectly. Reviewers should not assume that polished writing proves the analysis is accurate.
I use the term probability bias to describe a model's tendency to favor the most common pattern, even when the correct decision depends on an exception. That exception might be a local restriction or a technical dependency mentioned only once in a long document. Liu et al. (2024) found that models may give less attention to information located in the middle of lengthy inputs. The resulting summary can sound confident even when it omits the detail that controls the decision.
Decision-Grade Knowledge and Skill Development
Reading and writing are part of the work of thinking. Direct engagement with evidence exposes contradictions and forces the writer to decide what can actually be supported. When AI enters before that work, the user can receive a finished document without developing firm command of it.
Risko and Gilbert (2016) call the transfer of mental work to an external tool "cognitive offloading." The practice is routine and often useful. Its effect becomes visible when someone is questioned about a polished submission. The terminology may be correct, yet the writer cannot explain why one source controls or what evidence would change the recommendation.
In a survey of knowledge workers, Lee et al. (2025) found lower reported cognitive effort in AI-assisted tasks. That result says little about the quality of any one deliverable, but it matters when the person named as author cannot reconstruct the reasoning. Messeri and Crockett (2024) describe this gap as an illusion of understanding.
Table 1. Fluent AI Output and Decision-Grade Human Knowledge
| Fluent AI Output May Provide |
Decision-Grade Work Also Requires |
| A readable summary |
The controlling sources and any uncertainty the summary could not resolve |
| A polished recommendation |
A defensible rationale and an account of what would change the recommendation |
| A drafted technical requirement |
The operational need translated into testable acceptance criteria, with ownership identified |
| A concise executive briefing |
A clear decision request and the person authorized to accept the remaining risk |
Students and new professionals cannot rely on an established record. Their credibility depends on whether they can understand a difficult problem and explain the recommendation. AI can raise the apparent quality of the product faster than the author develops those abilities.
A seasoned professional may also become dependent if they begin tasks with a generated summary, which reduces the frequency of close reading and independent analysis. Gerlich (2025) reports an association between frequent AI use, cognitive offloading, and lower critical-thinking measures. The study is correlational, and its relevance here lies in the habit it describes: repeated offloading means less practice of the underlying skill.
Figure 1. Reported Correlations Among AI Tool Use, Cognitive Offloading, and Critical Thinking (Derived from Gerlich, 2025).
Operational Risk Across Contexts
In operational settings, a fluent document may move through review before anyone has examined the underlying problem in full. Mission and enterprise work depend on local knowledge that is difficult to reconstruct after it falls out of the record. When a summary is separated from its sources and decision logic, reviewers lose the ability to challenge the recommendation while correction is still practical.
Draft quality affects how quickly the gap is noticed: rough prose prompts questions, while a fluent draft can carry an unsupported assumption farther into review because it already sounds settled. Table 2 shows how that weakness plays out across three environments and which control fits each setting.
Table 2. AI-Assisted Work Risks Across Universities, Enterprises, and Defense Organizations
| Focus |
University Learning |
Enterprise Delivery |
Defense Mission Work |
| Core Risk |
The assignment reads beyond the student's actual command. |
The deliverable moves before the team tests the local constraint. |
The recommendation reaches staffing before mission validation. |
| What Is Lost |
The reading and reasoning the assignment was designed to assess. |
Business-rule context and clear ownership. |
Source authority and classification context; local mission knowledge. |
| The Result |
A grade becomes weak evidence of competence. |
Rework or a missed dependency appears late. |
An approver sees confidence before mission fitness. |
| Skill to Preserve |
Close reading and verbal defense. |
Requirements analysis across stakeholders. |
Ground-truth mission analysis and accountable approval. |
| The Control |
Assess the submission and the student's independent explanation. |
Require the team to identify the source basis and the tradeoff it accepted. |
Record the reviewer and the person accepting risk. |
DoD Adoption and Accountability
The Department of Defense has valid uses for generative AI, including requirements development and mission analysis. The Chief Digital and Artificial Intelligence Office (2025) identifies GenAI.mil as a Department platform supporting this work: any product entering the staffing process still needs a responsible person who can explain and defend it. Human oversight appears in the Department's Responsible AI Strategy and continues through its data and AI adoption strategy (U.S. Department of Defense, 2022, 2023).
Review begins before prompting. The team defines the decision and identifies the controlling documents. During drafting, reviewers look for omitted evidence and conflicting guidance. The National Institute of Standards and Technology (2023, 2024) supplies the broader risk framework, and the Office of Management and Budget (2025) sets the federal governance context in Memorandum M-25-21. The program office must still assign the subject-matter expert who verifies the product and the approver who accepts its use.
Deliverables should be considered untrusted until reviewers can trace their main claims to an inspectable source. A requirement may trace to an RFP section, while a project risk may come from an engineering log or meeting record. Sensitive work should rely on approved source documents. If no one can verify a claim's source, the document should not enter staffing.
Decision logic should be visible before AI improves the prose. A short rationale can record the team's position and the evidence behind it. Any unresolved uncertainty should stay on the page. Senior reviewers should attack the weakest link. If the owner cannot defend the document, it returns for revision regardless of how polished it looks.
Human Readiness and Governed Use
Gartner (2025b) distinguishes AI readiness from human readiness. Licenses and usage show deployment. Human readiness appears in whether personnel can recognize weak output and continue the work when the tool is unavailable. Adoption statistics cannot demonstrate either capability.
Leaders should test source judgment in practice. Can personnel identify the controlling document and check generated text against the actual constraint? The responsible owner should also be able to explain the decision without reading from the draft.
Gartner (2025a) predicts that, through 2026, some organizations will reintroduce AI-free assessments to test critical thinking under pressure. A periodic no-AI exercise can expose dependence before it reaches a live decision.
Controls should become stricter as the consequence of error rises. AI may organize evidence or draft a work product. People still define the issue and resolve conflicts in the record. Risk acceptance remains human.
Practical Controls in Existing Workflows
These controls fit inside review and staffing processes already in use. A general instruction to "use AI responsibly" leaves the reviewer without a standard to apply. Blanket bans usually push the work out of sight and leave no record of how the tool was used.
Figure 2. Practical Controls for Accountable AI Use
Papagiannidis et al. (2025) treat responsible AI governance as an organizational system. Figure 2 applies that premise to review processes already in use. A traceability matrix establishes the source set. A short decision rationale captures the author's position before drafting begins. The approval gate records who accepted the remaining risk.
In enterprise and DoD work, the traceability matrix should point to the exact source passage. A document title alone is too coarse for review. The review should also include a red-team-style challenge step. In this context, red-team review does not mean adversarial criticism for its own sake; it means assigning someone to test the weakest assumption, look for evidence the draft made less visible, and ask whether the author can defend the recommendation under pressure. Requirements should carry the source into acceptance criteria and test evidence.
In a course, permission for AI use clarifies what the assignment is meant to assess. An oral defense tests whether the student can explain the work independently. When needed, I ask for the prompt and the first output. Tracked changes then show where the student's own reasoning entered the work. Assignments built around recent local material or conflicting sources make generic consensus prose much less useful.
Applying the Controls Through Alchemy SDLC™
Embedding Controls in Delivery
A policy is only effective when the process makes people follow it. In schools, this may happen through oral presentations, teacher reviews, or tracked changes. In large operational environments with many people involved, depending only on people to notice every problem can lead to mistakes and inconsistency.
To reduce bias and protect human judgment, these controls need to be built directly into the work process. AI governance should not exist as a separate checklist that people complete by hand. It should be part of the technical lifecycle from the beginning.
For software, data, cyber, workflow, and mission-capability delivery, Alchemy SDLC™ provides a concrete, real-world example of how to build these theoretical controls directly into an operational pipeline. It does not replace human judgment. Instead, it structures human-AI collaboration so that every generated requirement remains traceable, inspectable, and defensible from intake to deployment (Campbell, 2026).
Figure 3 does not replace Figure 2. Instead, it shows how the controls in Figure 2 start to work inside Alchemist AI Pro™. Source materials, SME input, audit results, human revisions, acceptance criteria, and final requirements all become artifacts that can be reviewed. Later phases of Alchemy SDLC™ continue this process through development, user acceptance testing, deployment, and long-term sustainment.
Figure 3. Alchemist AI Pro™ Alignment to Accountable AI Controls
"Source-before-summary" becomes part of the requirements process. "Human-position-first" means that subject-matter experts explain and clarify the need before development begins. Visible verification includes:
- audit reviews
- checks for unclear language
- ongoing human review
A requirement is ready only when it can be traced to the original need, tested against its acceptance criteria, and explained by the person responsible for the work.
Later phases apply the same controls during delivery without requiring readers to understand how the platform works internally. AI-assisted development should still include checkpoints where people review the output before work moves to the next stage. Verification should test the technical result against the original business intent and the system artifacts. Sustainment should preserve all artifacts which enable traceability back to the mission need, allowing the next release to begin with a complete record of the current system and how it has changed over time. Teams should not have to rely on scattered prompts or knowledge held by only a few people. In Alchemy SDLC™, these records are kept throughout development, testing, deployment, and sustainment. AI-assisted work should remain open to review from the original need through future changes and updates.
AI can speed up research, drafting, coding, and testing, but the work must stay connected to its sources and remain available for review. Clear human owners, reviewers, and approvers must remain responsible for key decisions. Responsible AI delivery depends on traceable inputs, well-developed requirements, visible review, approval checkpoints, and records that are kept over time. Alchemy SDLC™ provides an example of how these controls can become part of the delivery process instead of being handled through a separate governance checklist.
A 30-Day Pilot
A 30-day pilot is enough to test these controls without changing the organization's structure. Choose one recurring product, such as a proposal or technical requirements package, and use the review process the team already follows. Before the pilot begins, identify the main source materials and the person responsible for approving the work. At the end of the month, compare how long the work took and how much revision it required. Then ask the owner to explain the final decision without looking at the AI-generated draft. The pilot should answer one practical question: Did AI improve the product while the responsible person still understood and controlled the sources and reasoning? The project record should provide clear evidence for the answer.
Conclusion
Generative AI will likely become a normal part of classrooms, project teams, and defense organizations. The main concern is whether people can still understand and inspect the work after AI has helped create it. A polished summary or draft is only useful when the person responsible can show where the information came from, explain the reasoning behind it, and identify the decision it is meant to support.
Students and new professionals should see this as part of building credibility. Credibility comes from reading the available information, dealing with uncertainty, and being able to defend a conclusion when questioned. Experienced teams have the same responsibility, but the consequences are often greater. Requirements, risks, recommendations, and approvals should include their sources, review comments, acceptance criteria, and the name of the person who accepts the remaining risk.
Alchemy SDLC™ shows how this responsibility can be built into the development process. Clear writing alone does not move a project forward. Source materials, input from subject-matter experts, audit results, user acceptance testing, saved records, and change history should remain connected throughout the lifecycle. The process should also connect the final work to the human reasoning that approved it. Requirements, code, testing, and future updates should follow that human decision-making instead of relying only on a single prompt or polished AI draft.
AI should reduce unnecessary work and help teams find problems earlier, while human reasoning remains clear enough to review, correct, maintain, and defend.
References
Campbell, J. (2026). The Alchemy SDLC [Presentation]. AI Pro Holdings, Inc.
Bender, E. M., Gebru, T., McMillan-Major, A., & Mitchell, M. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922
Chief Digital and Artificial Intelligence Office. (2025, December 9). The War Department unleashes AI on new GenAI.mil platform. U.S. Department of Defense. https://www.ai.mil/Latest/News-Press/PR-View/Article/4355177/the-war-department-unleashes-ai-on-new-genaimil-platform/
Gartner. (2025a, October 21). Strategic predictions for 2026: How AI's underestimated impacts are reshaping the future. https://www.gartner.com/en/articles/strategic-predictions-for-2026
Gartner. (2025b, October 20). Gartner survey finds all IT work will involve AI by 2030; organizations must navigate AI readiness and human readiness to find, capture and sustain value. https://www.gartner.com/en/newsroom/press-releases/2025-10-20-gartner-survey-finds-all-it-work-will-involve-ai-by-2030-organizations-must-navigate-ai-readiness-and-human-readiness-to-find-capture-and-sustain-value
Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), Article 6. https://doi.org/10.3390/soc15010006
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. https://doi.org/10.1145/3571730
Lee, H.P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. https://doi.org/10.1145/3706598.3713778
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://aclanthology.org/2024.tacl-1.9/
Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature, 627, 49–58. https://doi.org/10.1038/s41586-024-07146-0
National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.100-1
National Institute of Standards and Technology. (2024). Artificial intelligence risk management framework: Generative artificial intelligence profile (NIST AI 600-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.600-1
Office of Management and Budget. (2025). Accelerating federal use of AI through innovation, governance, and public trust (Memorandum M-25-21). Executive Office of the President. https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-21-Accelerating-Federal-Use-of-AI-through-Innovation-Governance-and-Public-Trust.pdf
Papagiannidis, E., Enholm, I. M., Dremel, C., Mikalef, P., & Krogstie, J. (2025). Responsible artificial intelligence governance: A review and research framework. The Journal of Strategic Information Systems, 34(2), Article 101885. https://doi.org/10.1016/j.jsis.2024.101885
Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688. https://doi.org/10.1016/j.tics.2016.07.002
U.S. Department of Defense. (2022). Responsible artificial intelligence strategy and implementation pathway. https://media.defense.gov/2022/Jun/22/2003022604/-1/-1/0/Department-of-Defense-Responsible-Artificial-Intelligence-Strategy-and-Implementation-Pathway.PDF
U.S. Department of Defense. (2023). Data, analytics, and artificial intelligence adoption strategy. https://media.defense.gov/2023/nov/02/2003333300/-1/-1/1/dod_data_analytics_ai_adoption_strategy.pdf
© 2026 ACC3 International. Alchemist AI Pro™ and Alchemy SDLC™ are trademarks of their respective owners.