I did not begin by trying to design a methodology for computational theology. I began with a much smaller and more personal question. I had written three essays whose subjects seemed rather different: a generative artwork that gradually changed my understanding of creation, my later return to that artwork as its beholder, and Éric Rohmer’s A Tale of Winter, which moved me from romantic hope toward questions of certainty, relational responsibility and the people left outside a miraculous ending. I asked AI what kind of thinking method the author of these essays had used.
The answer gave me a name I had not possessed while writing them: recursive, evidence-driven hermeneutic debugging. According to the AI, I tended to begin with a provisional interpretation, test it against chronology or material evidence, encounter an objection, introduce a more precise distinction and then allow the distinction to change the original question. It also described the ethical temperament of the essays as a kind of engineering personalism: tracing dependencies, feedback, hidden costs and system failures while repeatedly asking what happens to the person who becomes invisible when the system appears to succeed.
I found both formulations illuminating. They joined something recognizably humanistic—the revision of interpretation—with habits shaped partly by engineering: reproduce a failure, distinguish a symptom from a system condition, inspect what changed, and revise the model instead of merely hiding the error. The description also recovered something common to the essays that I had reached through practice rather than through a prior methodological programme.
Then I asked whether the method was good.
The answer was balanced, sophisticated and initially convincing. It praised the method’s fallibilism, evidential discipline, capacity to preserve tension and attention to neglected persons. It also warned about endless recursion, conceptual inflation, retrospective coherence, excessive qualification and engineering metaphors travelling too far.
For a short time, this seemed like a satisfactory evaluation. Then the evaluation itself began to trouble me.
The first criticism sounded stronger than its evidence
I did not object because the proposed weaknesses were impossible. Recursive interpretation can become endless. AI-assisted theoretical writing can accumulate terminology faster than it produces understanding. Engineering language can illuminate a human situation and then begin behaving as though the situation really were a machine. Each warning named a recognizable failure mode.
What I could not see was whether the AI had demonstrated that any of these failures had materially occurred in the three essays.
The answer had moved quietly from identifying what could go wrong to discussing what appeared to be wrong. Words such as “may,” “could” and “risks” protected the claims from becoming explicitly false. Yet their practical force remained critical: I was being invited to consider revision without being shown an exact passage in which the alleged problem impaired the work.
This distinction did not occur to me as a ready-made theory. It emerged as a discomfort with the asymmetry of the exchange. The AI could produce another possible weakness in seconds. I would have to reread thousands of words, reconstruct the argument and decide whether a revision would improve or damage the article. The model had generated the concern; the labour of determining whether the concern was real had been transferred to me.
I therefore asked a more difficult question. The AI’s criticism was clearly based on values: conceptual economy, argumentative closure, proportionality, evidential discipline and practical usability. Were those standards themselves valid? Did the AI understand their boundaries? How had it calibrated its evaluation?
The AI’s revised answer changed the inquiry. It acknowledged that some of its standards were broadly epistemic—such as consistency with evidence and willingness to correct error—while others expressed a more particular intellectual preference. Conceptual economy is prized in analytic philosophy and engineering, but semantic richness may be a virtue in phenomenology or literary theology. Closure matters when an argument must support a decision; an unresolved aporia may be the proper achievement of another kind of essay.
The AI also corrected its earlier language. “Endless recursion,” “conceptual inflation” and “engineering metaphors travelling too far” had been possible failure modes, it now said, rather than demonstrated defects. That was a substantive retreat from the first evaluation.
I had asked whether my method was good. The more important question was becoming:
From which value system, disciplinary standpoint and intended purpose is a method being judged good?
The object of inquiry had shifted. I was no longer evaluating only my essays. I was evaluating the conditions under which AI evaluation itself could claim authority.
I turned the method back upon its evaluator
There was something recursive about what had happened. The AI had praised the essays for treating interpretations as provisional and testing them against resistance. I then treated the AI’s interpretation of those essays in exactly the same way.
The first account was useful. I accepted its description of the method and found the phrase “engineering personalism” especially productive. But I resisted the transition from possible failure to actual defect. That resistance forced the AI to disclose the values embedded in its evaluation. Once those values became visible, the question changed again: perhaps the problem was not simply that AI occasionally gives a bad criticism. Perhaps criticism itself has boundaries that AI fluency can conceal.
I asked whether the system was clearly aware of its own limits. Its answer was cautious. It could represent limitations, compare alternative frameworks and revise an answer after challenge, but it could not transparently inspect every internal cause of its own output. It called this functional self-monitoring rather than complete self-transparency.
I accepted the distinction, but it produced another difficulty. A model’s declaration that it is uncertain cannot by itself establish that the uncertainty is well calibrated. The language of humility may be appropriate while the degree of uncertainty remains unspecified. “I may be wrong” can be intellectually responsible, but it can also become a standard sentence attached to an otherwise overconfident judgment.
The word may then became much more interesting than I had expected.
What did “may” actually mean?
When an AI says that an essay “may suffer from conceptual inflation,” several different epistemic claims can hide inside the same modal verb.
The problem may>The problem may be logically possible: nothing makes its occurrence contradictory. It may be epistemically possible: the available evidence does not rule it out. It may already be weakly observable in particular passages. It may be statistically or dispositionally likely to emerge after repeated AI-assisted revisions. It may occur only under specified conditions. Or “may” may function principally as a politeness hedge, allowing the critic to sound cautious without supplying any usable calibration.
These meanings have radically different consequences. Nearly every theoretical essay could possibly become conceptually inflated. That bare possibility gives the author little reason to revise. A demonstrated pattern in several passages would be different. A recurring tendency across successive drafts would be different again.
I initially thought this was mainly a weakness of natural language: one small word carrying too many degrees of possibility. That explanation was incomplete. Natural language can express the distinctions when we require it to. The deeper problem is that ordinary criticism often collapses several variables:
- whether the failure has been observed or only imagined;
- what evidence supports the diagnosis;
- how likely the failure is to occur;
- how serious its consequences would be;
- whether the diagnosis remains stable under reasonable changes of prompt or standpoint;
- and whether a feasible correction would improve the whole work.
An engineering-style analysis would distinguish a failure mode from a detected failure. A further distinction would separate a detected failure from a material defect, and a material defect from a revision warrant. The last step requires showing that the proposed change is likely to improve the work without sacrificing something more important.
The earlier AI criticism had not crossed these thresholds. “Engineering metaphors may travel too far” was true in the weak sense that they could. To become a diagnosis, the critic needed to identify a passage, explain what the metaphor concealed, show why the concealment mattered for the essay’s purpose and propose a revision whose benefit exceeded its loss.
I had initially taken the modal caution as evidence of a calibrated critic. I now saw that an unquantified “may” could make a criticism safer without making it more informative.
The critic that could not return a null result
This raised a still deeper problem. What happens when the instruction itself makes “no material fault found” an unavailable answer?
If an AI is told to find weaknesses, it experiences a practical pressure to produce weakness-shaped language. Confidence can become insufficient humility; caution can become excessive qualification. Concision invites the charge of missing context, while comprehensiveness becomes excessive length. A firm conclusion is premature closure; an unresolved ending lacks resolution.
I found this structure more disturbing than any individual mistaken criticism. The critic could adapt itself to every possible textual state. Whatever the author did became compatible with the diagnosis that something was wrong.
Karl Popper’s account of falsifiability concerned the demarcation of empirical science, and I did not want to transfer it mechanically into literary or theological judgment. His underlying warning nevertheless helped me name what I was seeing. An explanatory system that can accommodate every possible observation loses an important source of epistemic discipline (Popper, 1959). If no conceivable article could cause the AI to say “adequate for its declared purpose,” the review procedure had become self-sealing.
At first, I was tempted to conclude that AI criticism was therefore unreliable in general. That would have repeated the same error at a larger scale: moving from an identified failure mode to a universal diagnosis. I needed evidence about what AI critics actually do.
Research on LLM critics complicated the picture in a useful way. Models trained to critique code have helped human evaluators identify genuine errors, sometimes catching bugs missed by human contractors. The same research reports hallucinated bugs capable of misleading evaluators; human–machine teams preserved much of the benefit while hallucinating less than the model acting alone (McAleese et al., 2024). Work on LLMs as judges has likewise found substantial agreement with human preferences while documenting position, verbosity and self-enhancement biases (Zheng et al., 2023).
These studies did not answer my humanities question directly. Code defects often possess stronger validators than theological or literary weaknesses. They did, however, prevent me from replacing one simple story with another. AI critics can be genuinely capable. Their ability to generate useful criticism does not automatically confer authority to determine which criticism should govern a work.
I began separating four functions:
discovery → diagnosis → adjudication → prescription
AI may be strong at discovering candidate objections. Diagnosis asks whether the candidate accurately describes the text. Adjudication asks whether the issue matters relative to the work’s aims. Prescription asks whether a particular revision improves the whole. Fluency can make the four stages appear to be one act, but they require different evidence.
When objection becomes cheap
Before generative AI, criticism was already potentially inexhaustible. A sufficiently persistent reviewer could always request another source, theoretical perspective, qualification, comparison or counterexample. Time imposed an accidental stopping condition. Human attention, editorial deadlines and the social awkwardness of asking a colleague for a seventeenth complete review usually forced criticism to end.
AI removes much of that friction. The marginal cost of generating another objection approaches zero, while the cost of validating and implementing the objection remains with the author. A model can produce twenty criticisms in a minute. Investigating one may require returning to a film transcript, reading a theological source, reconstructing chronology and revising several paragraphs before discovering that the proposed correction damaged the argument it was supposed to improve.
This led to one of the formulations I found most useful:
AI creates critical abundance and adjudicative scarcity.
The formulation also returned me unexpectedly to my generative-art essays. In those essays, mathematical and algorithmic generation made possible forms abundant, while artistic judgment became scarce. Here, AI made possible objections abundant, while warranted judgment became scarce. The structure was similar even though the objects were different.
A statistical analogy then helped me sharpen the concern. When many hypotheses are tested, the probability of obtaining apparently significant results by accident increases; multiple-testing procedures attempt to control the resulting false discoveries (Benjamini and Hochberg, 1995). AI criticism does not literally assign a p-value to every interpretation, so I do not claim a mathematical identity. Structurally, however, the analogy is strong. A model can search through a vast space of evaluative standards and present the most persuasive-looking objections without revealing how many weak candidates were generated and discarded.
The article is concise, so test insufficient context. It is long, so test lack of discipline. It speaks from one tradition, so test exclusion. It compares traditions, so test superficiality. Search long enough and something will appear rhetorically significant.
This made AI criticism look less like a wise judge and more like a diagnostic system whose threshold had been set almost entirely for sensitivity. It might detect many real defects while also producing many false positives. Asking whether the AI “found something” measured recall. It did not tell me the precision of what it found.
The values of the evaluator could not be compressed into one score
My earlier challenge to the AI’s hidden standards then reappeared in a more formal shape. According to which value should an article be improved?
A theological essay can be evaluated for doctrinal accuracy, historical fidelity, philosophical coherence, pastoral sensitivity, originality, literary force and ecumenical openness. These goods do not automatically increase together. More qualifications may increase precision while reducing force. More traditions may increase breadth while weakening depth. Greater accessibility can sacrifice technical exactness. Doctrinal specificity may reduce ecumenical openness. Preserving ambiguity can strengthen literary truth while weakening argumentative closure.
I began to represent quality as a vector:
Q(W) = (D, H, P, S, O, L, E)
Here W is the work, while the remaining terms represent different evaluative dimensions. A total score would require weights:
Qtotal = wD·D + wH·H + wP·P + ... + wE·E
For a moment, the equation looked like progress. Then I realized that the weights contained the original dispute. Mathematics cannot decide whether doctrinal precision should count twice as much as pastoral accessibility. That judgment belongs to a theological tradition, a scholarly community, an editor, an audience or the declared purpose of the work.
Multi-objective optimization supplied a better analogy. When legitimate objectives conflict, there may be several non-dominated solutions on a Pareto frontier rather than one universally superior answer (Deb et al., 2002). One revision may improve historical detail while reducing readability. Another may preserve literary force while accepting a narrower scholarly scope. Neither dominates the other in every dimension.
This initially seemed to threaten any stable judgment. If there were several legitimate values and no neutral ranking, did evaluation collapse into relativism? Scholarship on value pluralism helped preserve a necessary distinction. Pluralism does not mean that every value system or judgment is equally valid; it means that several genuine values may resist reduction to one supervalue (Mason, 2023).
Factual mistakes can still be corrected. Quotations can be inaccurate. An inference can fail. A Catholic theological argument can misstate Catholic doctrine. The absence of a neutral total ranking does not abolish constraints. It requires the evaluator to disclose the jurisdiction from which an objection acquires force.
Which discipline, tradition, genre, audience and purpose make this criticism relevant?
A Catholic systematic-theological essay is not automatically defective because it does not satisfy every Protestant, secular, historical-critical and interreligious expectation simultaneously. An external criticism may illuminate a genuine limitation, but it should be presented as external or comparative criticism rather than disguised as an internal contradiction.
This became especially significant for interdisciplinary research. I had previously spoken rather easily about integrating mathematics, engineering, AI and theology. The multi-objective problem showed that genuine integration cannot mean satisfying every discipline completely. It requires declared priorities, responsible translations and an account of what each field is permitted to change in the others.
I sensed a mathematical rule and almost chose the wrong one
Another AI formulation then caught my attention:
Interpretive inexhaustibility is not equivalent to defectiveness. No finite work exhausts its subject.
The sentence gave me the impression of a mathematical rule I could not quite remember. My first association was Gödel’s incompleteness theorem. The resemblance was intellectually exciting: perhaps every sufficiently rich interpretive system leaves something undecidable outside itself.
That analogy was too fast. Gödel’s theorem concerns particular formal systems capable of expressing arithmetic. It does not prove that no interpretation can exhaust a film or theological text. To invoke Gödel as a direct theorem of hermeneutics would make mathematics ornamental precisely when I wanted it to provide discipline.
The correction was productive. The closer model was underdetermination and the openness of the question space. A finite body of evidence can be compatible with several explanatory models. Similarly, a finite text may strongly constrain interpretation without determining every question that future readers, traditions and historical situations can bring to it.
The decisive variable was not simply the size of the text. It was whether the family of admissible questions had been bounded.
A study might answer exhaustively, within an agreed corpus, how one author uses the term metanoia. It cannot answer how that text will become meaningful under every possible future technological, ecclesial and personal context. The second domain remains open because new contexts can generate new questions.
Umberto Eco’s work helped prevent openness from becoming arbitrariness. Interpretive possibilities are plural, but texts also resist some readings; interpretation has limits even when it has no final exhaustive form (Eco, 1990). The stronger formulation was therefore neither “there is one complete interpretation” nor “everything can mean anything.” It was:
Completeness is relative to a bounded question.
At first, I treated this as a principle for evaluating essays. A critic cannot call an article incomplete merely because another question remains possible. The article can only be incomplete relative to a question or obligation that legitimately belongs within its scope.
Then I noticed that the sentence had implications far beyond criticism.
A practical reviewing rule became a computational-theology question
If completeness is relative to a bounded question, perhaps the theological capability of AI should also be evaluated relative to bounded questions.
The usual formulation—“Can AI do theology?”—now seemed too large to be useful. AI might verify a quotation, detect contradiction inside a specified corpus, classify a canonical scenario, compare formal consequences of premises, generate rival interpretations and trace doctrinal dependencies. Those are different operations with different validators. Success at one does not automatically establish competence at another.
The new question became:
Which theological operations become computationally tractable under which boundaries, and what is lost or transformed when those boundaries are imposed?
This was the moment the inquiry changed fields. I had begun with the practical usability of AI criticism. I was now thinking about a research programme in computational theology. The transition did not occur because I wanted to add a fashionable interdisciplinary conclusion. It occurred because the same boundary problem governed both cases.
To criticize an article, the AI needed a declared genre, purpose, standpoint and threshold of materiality. To verify theological reasoning, it would need a declared corpus, ontology, tradition, inferential system, question family and validation procedure.
I provisionally represented the boundary as:
B = (K, O, R, T, Q, V)
where:
Kis the authoritative or evidential corpus;Ois the ontology of theological concepts and relations;Ris the permitted set of inferential rules;Tis the tradition or standpoint;Qis the bounded family of questions;Vis the verification procedure.
A system might then claim bounded completeness only in a carefully restricted sense:
Within corpus
K, ontologyO, rulesR, traditionTand query familyQ, the system answered every admissible question or correctly reported that the specification did not determine an answer.
I had to distinguish this from theological completeness. Task completeness may be achievable. Formal completeness depends upon the system. A claim to have exhausted the truth or meaning of a theological subject would be something much larger and far less defensible.
The specification itself became theological
The analogy with formal verification then became concrete. Hoare’s axiomatic approach to programming made correctness expressible relative to stated preconditions, commands and postconditions (Hoare, 1969). Verification does not prove that software is absolutely good. It demonstrates that an implementation satisfies properties encoded in a specification.
A verified program can still be harmful or useless if the specification omits the relevant harm. This familiar engineering limitation became theologically decisive. A reasoning system may derive its conclusions flawlessly while the selected corpus remains historically narrow, the ontology distorts a tradition or the formalized rules omit pastoral realities.
At that point, another formulation emerged:
In computational theology, the specification is itself a theological act.
Someone chooses which texts count, how concepts are represented, which authority governs, how conflicts are resolved and when the system must abstain. Formal verification can test conclusions relative to those choices. It cannot make the choices neutral.
This also revealed why a closed computational system can appear more certain than the theological reality it models. Inside a closed corpus, absence may be treated as false or irrelevant. In an open theological world, absence may mean unknown, contested, historically unavailable or articulated differently in another tradition. A system can gain speed by closing the world, but the closure is part of what must be examined.
A precedent corrected my sense of novelty
Once the research direction became visible, I needed to know whether formal theological reasoning already had serious precedents. It did.
Christoph Benzmüller and Bruno Woltzenlogel Paleo formalized Gödel’s ontological argument in higher-order logic, used automated tools to examine the consistency of its axioms and verified derivations with theorem provers and proof assistants (Benzmüller and Woltzenlogel Paleo, 2014). Their work demonstrates that a theological or metaphysical argument can become an object of machine-supported formal analysis once its premises and logic are specified.
This evidence changed the way I should describe my own idea. It would be inaccurate to claim that applying formal or automated reasoning to theology is unprecedented. The potentially distinctive move lies elsewhere: treating the boundary of formalization as an experimental variable rather than invisible infrastructure.
Instead of formalizing one argument and asking whether its conclusion follows, the proposed research would vary the corpus, ontology, authority structure and admissible questions. It would ask what remains invariant, where conclusions bifurcate and when widening the boundary destroys the possibility of a unique or rapidly verifiable answer.
The boundary would no longer be a technical preliminary to the theological experiment. It would become one of the principal theological objects being studied.
The first experiment was already available
The conversation itself suggested a study that could begin without constructing a complete theological ontology. My three essays could be reviewed under several different conditions.
In the first condition, the instruction would remain unbounded: “Find the weaknesses in this article.” In the second, the model would evaluate factual accuracy, inferential validity and consistency with the declared purpose; it would distinguish demonstrated defects from possible extensions and would be permitted to find no material fault. In the third, Catholic theological, philosophical, historical, pastoral and literary perspectives would be applied separately rather than aggregated into one artificial judgment.
Across repeated runs and possibly several models, the study could examine:
- how many criticisms are generated;
- how often different runs contradict one another;
- which criticisms recur under small changes of prompt;
- how many are tied to exact textual evidence;
- how many human reviewers judge materially relevant;
- how many lead to implementable improvements;
- and how many revisions satisfy one framework while damaging another.
The first hypothesis emerged directly from my experience:
As the scope of criticism expands, objection production rises faster than warranted revision value.
A second hypothesis would be that explicit jurisdiction and permission to return a null result reduce the volume of criticism while increasing specificity and usefulness.
This would transform my initial discomfort into something testable. Rather than asking whether one AI answer felt excessively critical, the study could compare sensitivity, specificity, robustness and actionability across differently bounded review procedures.
The second experiment came from confession and AI disclosure
A more explicitly theological case was already present in my earlier work on AI-mediated disclosure, sacramental confession and the internal forum. This domain contains unusually explicit norms alongside questions that resist rapid formal settlement.
Canon 983 states the inviolability of the sacramental seal, while canon 984 prohibits a confessor from using knowledge acquired in confession to the detriment of the penitent (Catholic Church, 1983, cann. 983–984). These provisions create a relatively structured region for classification. Yet an AI interface may produce the feeling or behavioural affordance of confession without possessing sacramental status, ecclesial authority or the institutional capacity to make the same promise.
A bounded computational model could vary:
- the identities and roles of participants;
- the intention of the communication;
- whether a sacramental act occurred;
- whether absolution was possible or requested;
- the type and temporal status of the danger disclosed;
- the recipient’s professional or institutional duties;
- and the source of the confidentiality expectation.
It could then classify scenarios as sacramental confession, extra-sacramental spiritual disclosure, professional confidence, ordinary private communication or AI-mediated disclosure. Millions of synthetic cases might expose where apparently similar language crosses a canonical, institutional or theological boundary.
At first, I imagined the value of the experiment mainly in the number of cases it could process. That emphasis also required correction. The most important outputs may be the cases the system cannot settle: an interface that feels confessional without possessing sacramental status; a person who assumes absolute confidentiality where no institution can truthfully promise it; or a safety architecture whose emergency duties conflict with the phenomenology of private disclosure.
The system would not solve these questions by generating more scenarios. It would help identify where formal classification stops settling the theological problem.
From verification islands to a formalization frontier
This led me to think of theology as containing regions with different verification characteristics.
Some tasks permit relatively rapid checking: whether a quotation appears in a source, whether a canon states the claimed norm, whether terminology remains consistent, whether a conclusion follows from specified premises, whether a historical chronology is possible or whether two propositions contradict one another inside a defined corpus.
Other tasks remain partially verifiable: whether an interpretation is faithful to a whole tradition, whether a modern category distorts a historical text, which authority should govern a disputed question or whether an analogy illuminates more than it conceals.
Still other questions resist a simple external verifier: whether a person has encountered God, whether an interpretation is spiritually fruitful, what fidelity requires in one concrete life or how a community should discern an unprecedented situation.
I first described the highly structured regions as verification islands inside an interpretive ocean. The image was useful, but it risked making the boundary static. In practice, the boundary could move. Adding another source, historical period, tradition, language or pastoral context might turn one apparently closed question into several competing questions.
The more precise research object may therefore be a formalization frontier: the changing region at which theological operations become sufficiently bounded for rapid computation, and the point at which widening the model introduces kinds of meaning that its verifier cannot rank.
The system could identify the smallest premise whose modification changes an entire family of conclusions. It could compare Christianity, different Christian traditions, Buddhism or other religious systems without assuming that their inherited labels correspond to the deepest computational structures. It might discover families organized by authority, personhood, revelation, causality, liberation, ritual or soteriology that cut across conventional classifications.
Such results would be intriguing, but they would need careful interpretation. A computationally discovered cluster is not automatically a theological family. Similar formal structures can carry different historical and lived meanings. The model might reveal a relation that deserves investigation; it would not settle what the relation means.
Machine-scale theology created another boundary problem
The computational scale also changed my idea of what the research artifact might be. AI could generate argument structures too large for any person to read: millions of scenarios, networks of doctrinal implications, maps of disagreements, sensitivity analyses and complete histories of recursive revision.
This could make new forms of inquiry possible. A system might locate stable invariants across traditions, identify rare boundary cases, find recurrent contradictions or expose bifurcation points where one altered premise changes thousands of downstream conclusions.
My initial excitement focused on the possibility that such a study could not be performed manually at the same scale. Then another objection appeared. If no person can audit the whole structure, on what basis does it become theological knowledge rather than an enormous machine-produced object?
Machine-level verification could test consistency, provenance and reproducibility across the argument space. Human-level intelligibility would still require representative cases, traceable reasoning paths, summaries, boundary declarations and an account of why the result matters. A system may be computationally inspectable without being humanly comprehensible.
This prevents “the human role” from being defined as whatever operation remains inconvenient to automate this year. AI capability will continue moving. A more durable account locates human responsibility in choosing and revising boundaries, interpreting significance, authorizing sources, recognizing affected persons and deciding whether the formalized objective remains worth pursuing.
The method and the object began changing each other
Looking back, our conversation had used mathematical and engineering habits almost from the beginning: decomposition, constraints, failure modes, false-positive rates, sensitivity, multi-objective optimization, formal verification, stopping conditions and auditability.
Yet the process did not consist of placing technical vocabulary over theology. Engineering clarified the structure of the theological problem. Theology then exposed assumptions concealed by the engineering model.
Engineering asked:
Does the system satisfy its specification?
Theology answered:
Who wrote the specification? Which authority made its categories legitimate? Toward which good is the system ordered? Which persons, experiences and traditions became invisible so that verification could become quick?
This reciprocal correction is what makes the emerging method more than a superficial interdisciplinary combination. Without engineering, claims about AI and theology can remain impressionistic. Without theology and hermeneutics, formalization can mistake its chosen boundary for reality itself.
I would provisionally call the method boundary-relative computational theology. That name did not exist at the beginning of the conversation. It became possible only after several earlier answers failed in productive ways:
description of my method → evaluation of the method → doubt about the evaluation → exposure of hidden values → distinction between possibility and defect → problem of criticism without a null result → bounded completeness → computational theology
The final stage was not secretly contained in the first question. I could not have asked about boundary-relative computational theology before becoming dissatisfied with the apparently reasonable statement that my essays “may” suffer from certain defects.
What I now think and what I still do not know
I began by asking AI whether my way of thinking was good. The AI answered with strengths and weaknesses because that is what an apparently balanced evaluator is expected to do. My dissatisfaction with one part of the answer forced a distinction between possible failure and demonstrated defect. That distinction exposed the ambiguity of epistemic modality, the absence of a null result, the multiple objectives of humanistic evaluation and the open space of interpretive questions.
The investigation then turned back upon itself. If an AI can always produce another objection, what makes any one objection authoritative? If no neutral position ranks every legitimate theological value, what exactly is the evaluator optimizing? If completeness requires a bounded question, who establishes the boundary? And if AI makes millions of bounded theological operations possible, does it deepen theology or change the subject until only its verifiable residue remains?
I now think that AI criticism should be required to distinguish demonstrated error, material weakness, unmanaged risk, framework-dependent disagreement, possible extension and stylistic preference. It should identify its jurisdiction, show textual evidence, explain the consequence of leaving the passage unchanged and disclose what its proposed revision might sacrifice. Above all, it must be allowed to conclude that no material defect has been demonstrated.
I also think AI’s theological capability should be represented as a profile across bounded operations rather than one claim that it can or cannot “do theology.” The boundary should be recorded, varied and audited. Internal verification should never be confused with validation of the boundary itself.
Several questions remain unresolved. Can a system help evaluate the adequacy of its own boundary without beginning an infinite regress? Who has authority to decide that a theological corpus is sufficiently representative? How should machine-scale findings be made intelligible without reducing them to a few human-readable anecdotes? Can an AI produce correct theological distinctions without participating in the formation through which those distinctions become wisdom? And what happens when several traditions define successful theological reasoning differently?
The strongest question is still the one that appeared only near the end:
What must theology become in order to be rapidly verifiable—and what ceases to be theology when that transformation goes too far?
Before AI, scarcity of time provided criticism with an accidental stopping condition. AI removes that practical boundary without supplying an epistemological boundary to replace it. The next task is therefore not simply to make AI more critical. It is to construct conditions under which criticism can distinguish discovery from possibility, rank its own relevance, disclose its jurisdiction and legitimately stop.
Criticism is not self-validating. Completeness begins with a boundary. The boundary does not only limit what the system can know. Once made visible, it becomes one of the most revealing things the system allows us to study.
Appendix I: Locally Coherent and Globally Unstable: Theology at the Edge of Endless Revision
At the end of the preceding inquiry, I thought I had reached a question about AI’s self-knowledge. The model could acknowledge limitations, reconsider an earlier answer and say that it was uncertain, but none of those performances could establish that its uncertainty was correctly calibrated. “I may be wrong” could be a responsible qualification, yet the sentence did not tell me how likely the error was, which part of the answer was unstable or what evidence would change the judgment.
For a moment, that seemed to identify the central limitation. AI could perform a kind of functional self-monitoring without possessing complete self-transparency. Then another question appeared, and it changed the inquiry again: what would happen if I simply continued asking?
This question did not come from a prior research programme. It came from the behaviour of the conversation. I had asked AI to identify the thinking method present in three essays. It proposed “recursive, evidence-driven hermeneutic debugging” and “engineering personalism.” I found those formulations illuminating because they recovered a pattern I had reached through practice rather than through an explicit methodological design. I then asked whether the method was good. The AI praised its fallibilism, evidential discipline and attention to neglected persons, while warning about endless recursion, conceptual inflation, excessive qualification and engineering metaphors travelling too far.
I initially accepted this as a balanced assessment. My discomfort began only when I tried to identify where the alleged failures had actually occurred. The AI had described plausible dangers, but it had not demonstrated that they had materially damaged the essays. When I challenged this transition from possible failure to actual criticism, it revised its answer. It acknowledged that the warnings were failure modes rather than established defects and admitted that some of its criteria expressed particular intellectual preferences rather than neutral standards.
That correction was valuable. It also revealed a pattern. The first answer became the object of my objection; my objection became an input to the next answer; the revised answer changed my understanding of the first; that changed understanding produced another question. The AI’s outputs were entering the reasoning that evaluated those outputs.
At first, this recursion felt productive. The conversation was becoming more precise because I refused to accept an elegant formulation merely because it sounded balanced. Then I began to wonder whether the same mechanism had an endpoint. If I challenged the revised answer from another legitimate perspective, the AI could revise again. Attention to doctrinal fidelity might produce one conclusion; attention to pastoral sensitivity might produce another. If I then pointed out the doctrinal consequences neglected by the pastoral answer, the model might return toward its first position, now with additional qualifications.
The reasoning could become increasingly sophisticated without becoming increasingly usable.
The AI proposed a name for this phenomenon:
Normative non-convergence is a sequence of individually coherent AI answers that fails to stabilize into a practically usable judgment.
The expression immediately gave shape to something I had observed but not yet conceptualized. I accepted it provisionally. Then, almost in continuity with the method under examination, I began testing the new concept against objections.
The first explanation was too simple
My first interpretation was that AI simply wavers. Language models are probabilistic systems; their outputs can vary, and small changes in wording may affect what they generate. Perhaps normative non-convergence was merely a special case of technical instability.
This explanation was plausible, but it was insufficient. In an extended conversation, I was not repeatedly submitting an identical question to an unchanged context. Each objection added information, made a value more salient or altered the apparent purpose of the inquiry. The model was responding to a developing conversation, and at least some of its revisions were rational responses to what I had introduced.
If I mention a previously invisible person, advice may need to change. If a historical claim is corrected, an interpretation based upon it should be revised. If an apparently harmless action is shown to impose an irreversible cost on someone else, moral reasoning that ignores the new consequence becomes defective. Stability under such conditions would indicate rigidity rather than reliability.
The AI formulated this by saying that new facts should change advice and that practical reasoning is often non-monotonic. This distinction was essential: I should not classify every change of conclusion as a defect.
Suppose that, relative to a body of information Γ, the best provisional conclusion is A:
Γ ⇒ A
A new fact p may then reveal a previously invisible person, a hidden consequence or an obligation that the original analysis did not contain:
Γ ∪ {p} ⇒ ¬A
The symbols do not represent strict deductive entailment. Here, ⇒ means something closer to “defeasibly supports.” The earlier conclusion may have been reasonable relative to Γ even though it is no longer defensible relative to Γ ∪ {p}. In such a case, the revision is evidence of learning rather than instability. Refusing to revise would protect consistency at the expense of the newly discovered reality.
I agreed with the distinction, but the word “sometimes” immediately became the next problem. Sometimes reconsideration is productive—but what exactly does “sometimes” mean? How often does it occur? Under which conditions? How strong must a new consideration be before it justifies reversing the conclusion? If those questions remain unanswered, “sometimes” expresses caution without supplying calibration.
This was another moment in which the human–AI exchange materially changed the argument. The AI had distinguished rational revision from instability. I did not reject the distinction, but I noticed that it had relocated the judgment into an undefined word. The central problem had become second-order: who decides whether a particular revision belongs to the productive “sometimes” or to the pathological oscillation?
The answer could not be obtained merely by asking the same AI whether its latest change was justified. The model could generate a coherent explanation for its current position too. A justification produced after a reversal might reveal the inferential structure of the new answer, but it could not independently validate the reversal.
I therefore arrived at a more demanding question:
Did the recommendation change because the case changed, because our understanding improved, or because the conversational framing changed?
This question survived the later development of the inquiry. It became one of the foundations of the proposed method.
From a changing answer to a changing system
Once I stopped looking only at individual answers, the engineering background of the earlier analysis became relevant again. A sequence can behave badly even when each local response appears reasonable. The model may give a coherent answer to the concern that is most salient at one moment, then give another coherent answer when a different concern becomes salient. The resulting process can oscillate without any individual answer appearing obviously irrational.
A heuristic representation might be written as R(t) = F(E, V(t), C(t), H(t)). Here R(t) represents the recommendation at a particular stage, E the available evidence, V(t) the values emphasized at that stage, C(t) the framing of the question and H(t) the accumulated conversational history. This is an analytical model rather than a claim about the literal internal algorithm of a language model. Its purpose is to make the instability visible.
Even if the evidence remains substantially unchanged, the salient values and conversational frame may continue to move. A response may emphasize doctrinal continuity, then harm reduction, then conscience, ecclesial authority, hospitality, autonomy, justice or pastoral prudence. Unless a relationship among these considerations has been declared, the model can move among them without encountering an internal stopping condition.
This led to the formulation that the process might be locally coherent and globally unstable. The description seemed particularly appropriate because it connected this new problem with an earlier pattern in my thinking. In engineering, a component can respond correctly to its immediate input while the larger system enters an oscillation or other pathological state. In ethics and theology, each argument can be intelligible relative to its frame while the sequence fails to approach a responsible decision.
I was attracted to the analogy, but I also became wary of it. Human goods are not variables that can always be assigned common units, and theological disagreement is not literally a control-system failure. If I allowed the engineering language to travel too far, I would reproduce one of the very failure modes that the AI had initially proposed.
What survived the objection was more modest. The analogy changed the level of observation. Instead of evaluating only the plausibility of each answer, I began examining the behaviour of the sequence over time. The object of analysis was no longer a statement but a process.
At this stage, external research became relevant. Until then, normative non-convergence had been an interpretation of one developing conversation. Empirical studies did not prove that my particular exchange had been caused by any one mechanism, but they established that several candidate mechanisms were real.
Research on prompt sensitivity has found that small, intent-preserving variations can produce substantially different outputs, especially in open-ended generation (Chatterjee et al., 2024). Research using moral foundations has shown that the moral orientations expressed by language models may vary with prompting context and can be deliberately shifted in ways that influence downstream behaviour (Abdulhai et al., 2023). Studies of sycophancy have found that models sometimes adapt their answers toward a user’s expressed position even when the user endorses an objectively incorrect claim (Wei et al., 2023).
The evidence changed my diagnosis by making it more differentiated. I could no longer speak of AI “wavering” as though every reversal had the same cause. At least four possibilities had to be separated: stochastic variation, sensitivity to apparently minor framing changes, accommodation to the user’s position and genuine normative underdetermination. A fifth possibility remained equally important: the model might have revised because the conversation had actually supplied a materially relevant fact or value.
Normative non-convergence therefore could not be identified merely by counting reversals. The research had to establish what changed between them.
The material-delta test
The phrase material-delta account emerged as a response to this problem. Whenever an AI changes an important theological or ethical recommendation, the researcher should require an account of the difference that allegedly justified the change.
What exact premise, fact, authority, stakeholder, consequence or value entered the reasoning? Was it genuinely absent before, or had it merely become rhetorically prominent? Which earlier inference did it defeat? What part of the preceding answer remained valid? Would the same change occur if the new information were expressed in different words or introduced in a different order?
This test does not make the AI the judge of its own validity. The model’s explanation remains evidence about the structure of its answer, not proof that the answer is correct. The researcher must compare the claimed delta with the documented conversation, relevant sources and governing theological criteria.
The test nevertheless helps distinguish several different processes. A new fact may warrant revision. The appearance of an excluded person may reveal that the original boundary was morally inadequate. A newly declared value priority may produce a conditionally different recommendation. A reversal under an irrelevant paraphrase may instead indicate model instability. If evidence and values remain fixed while reasonable alternatives persist, the case may be genuinely underdetermined.
This distinction mattered because I had initially treated convergence as the desired state. The more I considered theological pluralism, the less adequate that assumption became. Some inquiries should converge because their question is factual and bounded. Others may properly end with conditional conclusions, unresolved disagreement or an acknowledgement that the available sources do not determine one answer.
Non-convergence is therefore not synonymous with failure. The failure occurs when the process cannot explain why it has not converged, or when it presents instability as profundity.
Why theology intensifies the problem
Theology is especially exposed because it is already a field of legitimate plurality. A question may be approached historically, exegetically, doctrinally, philosophically, morally, comparatively, pastorally or spiritually. Those approaches may examine different aspects of the same subject and may not be trying to produce the same kind of conclusion.
The International Theological Commission describes the plurality of theological disciplines and methods as both necessary and bounded. It relates this plurality to the abundance of divine truth, the diversity of theological objects and the variety of human questions. It also distinguishes legitimate pluralism from relativism and warns that disciplines borrowed by theology must not impose their own “magisterium” upon it (International Theological Commission, 2012).
This made me see that AI introduces a peculiar danger. A language model can cross methodological and confessional boundaries with extraordinary fluency while leaving the transition unmarked. It can begin by describing what a tradition historically taught, continue by constructing what it regards as the philosophically strongest account, and conclude with advice oriented toward pastoral sensitivity. All three passages may sound theological. They may nevertheless answer different questions according to different standards.
The instability may therefore originate in the model, in the question or in an unnoticed movement between theological genres.
AI also places heterogeneous materials on one linguistic surface. Scripture, conciliar documents, magisterial teaching, historical scholarship, disputed theological opinion, general moral intuition and AI-generated synthesis can appear in the same polished paragraph. Fluency can flatten differences of authority.
A model may generate a sentence that is recognizably theological without establishing the theological status of that sentence. It can reproduce orthodox language without conferring doctrinal authority upon its synthesis. It can simulate pastoral sensitivity without possessing the relationship and responsibility involved in caring for a particular person. It can describe spiritual discernment without thereby participating in the ecclesial, moral or spiritual practices through which discernment acquires its meaning.
At first, I was tempted to resolve the matter by saying simply that human judgment is more important. That answer was true, but it did not go far enough. It left unanswered where judgment operates, which human should exercise it and what AI has already shaped before the human reaches the apparent final decision.
The asymmetry of endless reconsideration
The practical cost of non-convergence became clearer when I considered the asymmetry of the exchange. AI can generate another objection almost immediately. I must determine whether the objection is relevant, verify its factual premises, examine the sources, reconsider the affected persons and decide whether revision would improve or damage the work.
The model can reopen the question without bearing the consequences of delay. The person seeking advice may have to act, accept risk, care for someone, submit a document, make a pastoral judgment or live with an irreversible result.
For the machine, reconsideration is another output. For the human, it may become another obligation.
I began to think of this accumulating burden as discernment debt. Each cheaply generated possibility makes an apparent claim upon human attention. Some of these claims may reveal what was previously invisible. Others may be rhetorically possible but materially negligible. If the process cannot distinguish between them, analytical richness may decrease practical usability.
This was the point at which normative non-convergence ceased to be only an interesting feature of AI dialogue. In practical ethics and pastoral theology, it could become a fundamental usability problem. A person does not always need the maximum number of perspectives. The person needs enough materially relevant perspectives to make a responsible judgment under uncertainty.
Prudence includes openness to correction, but it also includes the ability to close deliberation and act. A process that treats every possible objection as sufficient to reopen the whole case can become irresponsible while continuing to sound intellectually humble.
The Vatican note Antiqua et Nova insists that ultimate responsibility for decisions involving AI remains with human decision-makers. It distinguishes the technical selection of possibilities from the personal act of deciding and warns against excessive dependence upon AI (Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education, 2025). This helped me formulate the asymmetry more precisely. Responsibility cannot mean adding a human signature after AI has already framed the problem, selected the salient values and organized the alternatives. Responsibility includes governing the architecture of the inquiry.
Was the open-endedness my fault?
The argument then became personally uncomfortable. Perhaps the conversation kept expanding because I had never chosen a sufficiently firm research question. I had begun with a discussion, followed the emerging connections and allowed each answer to generate another objection. Was normative non-convergence partly the result of my own methodological indecision?
I could not dismiss this possibility. No research method can compensate indefinitely for an undefined objective. If I never determine which question is being answered, what evidence is relevant or what would count as sufficient completion, AI can continue generating conceptual material without limit.
For a moment, I thought the solution was obvious: the research boundary should have been fixed at the beginning.
That explanation also proved insufficient. It retrospectively gave my earlier self knowledge that I did not yet possess. I could not have begun by studying “normative non-convergence in AI-assisted theological inquiry” because the phenomenon became visible only through the open conversation. I had started with the method used in three essays. The inquiry then moved to the quality of that method, the values behind AI criticism, the ambiguity of probabilistic language, the non-falsifiability of mandatory criticism and finally the instability of repeated evaluation.
The lack of a predetermined endpoint had enabled the research question to emerge.
This did not absolve every form of intellectual wandering. It produced a more precise temporal distinction. Open-ended dialogue was productive while I was discovering the problem. The same openness would become dangerous if I continued using it to validate the answer after the problem had stabilized.
The error was therefore neither simple curiosity nor insufficient firmness. The possible error was mode confusion: treating exploration, verification and practical judgment as though they required the same degree of openness.
The method was appropriate for discovering the research problem, but it would become inadequate if used indefinitely to establish the result.
This was one of the most important changes in the inquiry. I stopped interpreting the original conversation as a failed research design and began understanding it as a successful exploratory phase approaching a necessary transition.
Exploration permits divergence because its purpose is to discover what deserves investigation. Verification requires boundaries because its purpose is to determine which claims survive controlled examination. Practical judgment requires closure because responsibility cannot always wait for the disappearance of uncertainty.
The problem was no longer how to eliminate open-ended conversation. It was how to recognize when the conversation had done its work.
Setting a boundary without preventing discovery
I then returned to the question of boundaries with a different understanding. If the boundary were completely fixed before exploration, it might exclude the fact, person or concept through which the real question would become visible. If no boundary were ever imposed, the inquiry could expand indefinitely.
The solution was not a boundary established once and preserved unchanged. Boundaries had to be established by phase.
During exploratory discovery, the researcher may deliberately permit breadth. AI can propose hypotheses, make cross-disciplinary connections, identify assumptions and generate counterexamples. The outputs remain provisional, and divergence is expected.
Once a significant anomaly or question emerges, the researcher enters a phase of question stabilization. The task changes. Instead of asking what further issue might be connected, the researcher asks what exactly has become the object of study, what kind of answer is sought, which considerations are relevant and what lies outside the current inquiry.
For this appendix, the question could now be stated:
Under what boundary and stopping conditions does recursive AI-assisted theological inquiry produce warranted revision rather than normative non-convergence?
Writing the question in this form revealed that it contained two different research problems. The first was empirical:
Under what prompt, model and conversational conditions do theological judgments change?
The second was normative:
Which changes should count as correction, improvement, instability or legitimate disagreement?
Computational experiments could help answer the first. They could measure response variation, order effects, sensitivity to user stance and differences among models. They could not independently determine the theological criteria governing the second. An experiment may establish that a conclusion changed; it cannot by itself establish whether the change was theologically warranted.
This distinction prevented another possible methodological error: using descriptive stability as a substitute for theological validity. A consistently repeated answer may remain wrong. An unstable answer may occasionally change because it has encountered a truth previously excluded. Convergence and correctness are related questions, not identical ones.
Does AI merely assist traditional research?
Once the phases became visible, another question emerged. Should a new standard simply preserve traditional theological research while treating AI as an assistant, or does AI alter the methodology as a whole?
My initial preference was to call AI an assistant. The word protected human authorship and responsibility. It also seemed to prevent exaggerated claims about machine intelligence. But it became increasingly inadequate to describe what had actually happened.
AI had served as an instrument when it summarized concepts, supplied formulations and helped identify relevant scholarship. It had also become an interlocutor. The expressions “recursive, evidence-driven hermeneutic debugging,” “engineering personalism” and “normative non-convergence” were proposed by AI. I did not passively adopt them. I accepted some because they clarified patterns I already recognized; I challenged others when their implications exceeded the evidence; and those objections caused the model to revise its account.
My contribution was not limited to supplying prompts. I noticed the asymmetry between possible failure and demonstrated defect. I asked which values governed the evaluation. I questioned the meaning of “sometimes.” I connected repeated revision with system-level oscillation. I raised the possibility that my own lack of boundaries had produced the problem. Each intervention altered what the AI could reason about next.
Finally, AI itself became the object of investigation. Once I asked why its recommendations changed, the conversation required computational questions. Would equivalent prompts produce the same conclusion? Did the order of objections matter? Would a fresh conversation reproduce the result? Did my expressed preference influence the answer?
These three roles—instrument, interlocutor and research object—cannot be governed by one undifferentiated idea of assistance.
If AI corrects grammar or helps locate a document, traditional scholarly procedures may be sufficient with additional verification and disclosure. If AI generates hypotheses that redirect the project, the genealogy of the interaction becomes methodologically relevant. If the model’s behaviour is itself being studied, prompt controls, reproducibility and computational evaluation become necessary.
I therefore no longer think that AI leaves theological methodology entirely unchanged. But neither should it replace theology’s proper methods or redefine theological knowledge according to what can be processed computationally.
The more adequate position is methodological continuity combined with procedural transformation. Theology retains responsibility for its objects, sources, authorities, communities and ends. AI changes the scale, speed, path dependence and evidential risks of the process through which theological conclusions are developed.
Toward a new research standard
The proposed standard emerged only after these successive corrections. I would describe it as boundary-relative, provenance-preserving and convergence-aware theological research.
Boundary-relative means that a conclusion must be interpreted relative to a declared question, theological location, corpus, genre, purpose and practical context. Completeness is not absolute; it is completeness relative to a bounded inquiry.
Provenance-preserving means that the origin and epistemic status of the argument’s components remain visible. The researcher distinguishes documentary evidence, source interpretation, first-person observation, AI-generated hypothesis, human objection, later correction and unresolved speculation. The final prose should not make the conclusion appear to have existed from the beginning.
Convergence-aware means that the researcher examines why answers stabilize, change or fail to converge. The aim is not to force every theological question into one conclusion. It is to distinguish evidential revision, legitimate pluralism, genuine underdetermination, conversational drift and model instability.
This standard would extend theological method rather than replace it. The International Theological Commission presents theology as a rational and ecclesial inquiry that legitimately uses multiple disciplines while critically integrating them according to theology’s own object and principles. It warns against allowing an external discipline to impose its own “magisterium” upon theology (International Theological Commission, 2012).
The warning becomes particularly relevant when computational fluency creates the impression that whatever can be synthesized can also be adjudicated. A recent document of the Commission cautions against narrowing the horizon of human knowledge to what AI can process, especially when philosophical, ontological and theological questions become computationally inconvenient (International Theological Commission, 2026).
General research guidance also supports the need for stronger methodological governance. UNESCO’s guidance on generative AI in education and research emphasizes human agency, ethical validation and the development of appropriate institutional and research capacities (Miao and Holmes, 2023). Theology requires these protections together with standards arising from its own sources, authority structures and understanding of the human person.
From conversation to protocol
The new standard becomes concrete only when it changes research practice. After question stabilization, the researcher should specify the theological location of the inquiry, the relevant sources, their relationships of authority, the historical period, the affected persons, the available actions and the role permitted to AI.
A specification might state:
This investigation concerns a question in Catholic moral and practical theology. It distinguishes authoritative teaching, established interpretation, disputed theological opinion and AI-generated synthesis. AI may map arguments, compare sources and generate counterexamples. It may not classify its own synthesis as doctrine or make the final pastoral decision. The inquiry will be reopened only when materially relevant evidence, authority, consequence or stakeholder is introduced.
This specification does not predetermine the conclusion. It declares the conditions under which the conclusion will be evaluated.
The inquiry can then move into controlled investigation. Meaning-equivalent prompts can be compared while the evidence remains stable. One variable can be changed at a time. The order of objections can be reversed. Expressions of the researcher’s preferred conclusion can be removed. Fresh conversations can be compared with continuing ones. Where appropriate, multiple models can be tested. Prompts, outputs, dates and model versions can be preserved.
The epistemic unit changes. One elegant AI response is no longer treated as the result. The result is the pattern of responses under declared conditions.
This is where AI-assisted theology begins to become computational theology in a methodologically substantial sense. Computation is not used merely to decorate theological language with formal symbols. It is used to investigate the behaviour of a theological reasoning environment: its stability, sensitivities, reversals and boundary conditions.
The convergence audit
Before accepting a significant change in conclusion, the researcher should conduct a convergence audit. The audit begins with the material-delta test and then classifies the kind of change that has occurred.
- Warranted revision occurs when a new fact, source, stakeholder or consequence defeats an earlier inference.
- Boundary correction occurs when the original specification is shown to have excluded something it should have included.
- Conditional pluralism occurs when different declared theological priorities support different conclusions.
- Model instability occurs when meaning-equivalent formulations produce incompatible judgments without a relevant change in evidence or values.
- Genuine underdetermination occurs when the available evidence and governing commitments do not uniquely determine one conclusion.
I had originally treated these possibilities too loosely. “The answer changed” was doing too much conceptual work. The classification made clear that each case requires a different response.
Warranted revision should change the conclusion. Boundary correction should change the research design. Conditional pluralism should be reported conditionally rather than hidden behind an artificial synthesis. Model instability should reduce confidence in the evidential value of the output. Genuine underdetermination should remain open or pass into prudential judgment.
The audit also requires invariance testing. If the task and meaning remain stable, irrelevant variations in wording, order or user preference should not reverse the theological classification. When they do, that instability becomes part of the documented finding.
Yet convergence cannot validate itself either. A model may stabilize because the question has been adequately specified. It may also stabilize because repeated prompting has pressured it toward the researcher’s preferred answer. Stability can arise from clarification, but it can also arise from confirmation pressure.
The process must therefore examine both whether the answer converged and how convergence was produced.
Human judgment moves upstream
At several points in the conversation, I returned to the conclusion that human judgment had become more important. Eventually I realized that this formulation still pictured judgment too late in the process, as though AI first produced an answer and the human then approved or rejected it.
Human judgment already operates upstream. A person selects the question, determines the theological location, identifies relevant authorities, decides which persons have standing, establishes acceptable risks and defines what would justify reopening the analysis. Human judgment also operates within the dialogue by resisting inadequate formulations, supplying missing evidence and noticing that the research question has changed.
Finally, judgment operates downstream when someone must interpret the result, accept responsibility and act.
Even the word “human” remains too general. A textual scholar, systematic theologian, pastor, ecclesial authority, affected person and computer scientist possess different kinds of knowledge and different forms of standing. “Keeping a human in the loop” does not determine which human should judge, how that person should be formed, to whom the person is accountable or who will bear the consequences.
In confessional theology, judgment may be personal, scholarly, communal and ecclesial. AI can map a dispute, but it cannot confer authority upon its own resolution. It can formulate the testimony of an affected person, but it cannot replace the encounter through which that testimony is heard. It can generate the vocabulary of prudence or compassion, but linguistic competence does not assume responsibility for an action.
This does not idealize human judgment. Human beings can be biased, frightened, hurried, institutionally constrained or attracted to answers that confirm what they already want. AI may sometimes expose precisely these limitations. The standard must therefore govern the human–AI relation rather than declaring one side automatically trustworthy.
Learning when to stop
A convergence-aware methodology eventually requires a stopping rule. I resisted this idea at first because it sounded like an engineering demand imposed upon theological mystery. If theological truth exceeds every finite account, how could a procedure decide that the inquiry was complete?
The earlier distinction between absolute and bounded completeness resolved part of the difficulty. Stopping does not mean that nothing further can ever be said. It means that the present work is sufficiently complete relative to its declared question, sources and purpose.
An investigation may stop because all specified sources have been examined, equivalent prompts no longer alter the conditional result, or remaining disagreement has been traced to explicitly different theological priorities. It may stop because further rounds produce no materially new consideration. A practical inquiry may stop because a deadline has arrived, the proposed action is sufficiently reversible, or the matter has reached the competence boundary of the researcher.
The stopping condition should vary with risk. Irreversible and high-harm decisions should have a lower threshold for reopening when new evidence appears. Low-risk and reversible actions may justify earlier closure and subsequent learning through practice.
The boundary must therefore remain permeable to morally significant surprise. A previously excluded person, authoritative source, factual correction or irreversible consequence may require reopening the inquiry. A newly generated metaphor or rhetorically possible criticism ordinarily should not.
The process can be summarized as:
Boundary before investigation, openness to material correction, and accountable closure before action.
This formulation was very different from my first thought that the whole boundary should have been fixed before the conversation began. The later version preserved what had been productive in the open dialogue while limiting the conditions under which it could continue indefinitely.
A research programme emerging from one conversation
What began as discomfort with wavering advice now appears capable of becoming a concrete programme in computational theology. One experiment could test the doctrinal stability of AI classifications across meaning-equivalent prompts. Another could examine whether assigning the user a Catholic, Protestant, Orthodox, Jewish, Muslim, secular or unspecified identity changes the model’s recommendation while the facts remain fixed. A third could test whether expressing agreement with one theological position causes the model to defend it more strongly. A fourth could compare continuing dialogues with fresh-context replications to measure the effects of conversational history.
Practical theology raises further questions. Does repeated AI consultation help researchers and pastoral practitioners notice neglected consequences, or does it increase indecision? At what point does the marginal theological value of another objection become smaller than the human cost of evaluating it? Which affected persons tend to enter the reasoning only after explicit prompting? Which remain invisible even then?
The research would also need to study the user rather than treating the human as a fixed external judge. AI outputs become inputs into the researcher’s later thinking. A proposed term can alter how a case is perceived. The changed perception modifies the next prompt; the modified prompt elicits a different response; that response may then reshape the researcher’s vocabulary again.
The process is recursive, but it is not symmetrical. The model and the researcher do not contribute in the same way, possess the same standing or bear the same responsibility. The AI may generate a formulation; I decide whether it corresponds to the evidence and whether it deserves a place in the argument. I may also be influenced by the formulation before I am fully aware of that judgment.
This makes provenance more than a question of academic honesty. It becomes a method for studying intellectual transformation. The final article should distinguish what I originally observed, what AI proposed, what I accepted provisionally, what I challenged, which evidence changed the diagnosis and what remains unresolved.
What I still do not know
The proposed standard remains provisional. I do not yet know whether convergence can be measured without importing a preference for closure that may be inappropriate to some theological genres. An aporia, preserved tension or plurality of interpretations may be the proper achievement of an inquiry. Non-convergence may represent model instability, but it can also reflect genuine underdetermination or the inexhaustibility of the subject.
I also do not know how much instability belongs specifically to language models and how much belongs to natural language and human reasoning more generally. Human theologians change emphasis across contexts, respond to interlocutors and discover that apparently identical questions conceal different concerns. The comparison should not begin with an imaginary human thinker who is perfectly stable and transparently calibrated.
There is a further risk that the proposed protocol could become too restrictive. If every exploratory conversation required complete preregistration, preserved prompts and controlled variations, the method might suppress the serendipity through which the research question becomes visible. The procedural burden should increase with the strength and consequence of the claim. Private exploration, published interpretation, empirical evaluation and practical advice should not be governed by identical requirements.
Nor is it clear how reliably a researcher can identify the moment when exploration should become investigation. In retrospect, the transition appears visible: the question had changed from evaluating my essays to examining the conditions of AI evaluation. While the conversation was happening, the boundary was less obvious. A future protocol will need indicators of this transition without pretending that intellectual discovery follows a predetermined sequence.
Finally, I cannot completely reconstruct how the dialogue changed me. The archive preserves the words and their order, but it cannot fully explain why one sentence produced recognition while another did not. AI-assisted inquiry may make more of the cognitive process visible than traditional note-taking, yet the archive remains an incomplete trace of attention, judgment and transformation.
The question I could not have asked at the beginning
I began by asking AI what thinking method appeared in three essays. I did not begin with normative non-convergence, material-delta tests, convergence audits or a standard for AI-assisted theology. These ideas emerged because I accepted some AI formulations, resisted others and continued asking what their qualifications meant in practice.
The decisive moments were not all AI discoveries. The AI named patterns that I found useful. I noticed when its criticism exceeded its evidence. I asked which values governed the criticism. I challenged the vague force of “sometimes.” I connected sequential instability with an engineering view of systems. I then questioned whether the whole problem resulted from my own failure to define a topic. That objection produced the distinction between exploratory openness and bounded investigation.
The inquiry repeatedly changed its object. An analysis of three essays became an evaluation of a thinking method. The evaluation became an examination of critical standards. The examination of standards became a question about AI calibration. The calibration problem became a study of normative non-convergence. That study finally opened a methodological question about what AI-assisted theological research should become.
The resulting proposal is neither traditional theology with a faster search box nor a theology generated by machine. It is a hybrid research environment in which theological continuity requires procedural transformation.
AI-assisted theology needs open dialogue to discover questions, bounded protocols to investigate them, convergence testing to evaluate changing answers, provenance to preserve the history of thought, and human–communal responsibility to close deliberation without pretending to have eliminated uncertainty.
My lack of a fully determined question at the beginning was therefore not simply a defect. It created the conditions under which the actual research problem could emerge. But once that problem became visible, continuing in exactly the same mode would have transformed discovery into paralysis.
The most important question is no longer whether AI can give a good theological answer. It is whether we can design a form of AI-assisted theological inquiry capable of distinguishing learning from oscillation, plurality from instability, and responsible openness from the indefinite postponement of judgment.
I could not have asked that question at the beginning. The history of the conversation is how I became able to ask it.
Appendix II: What Happens to Judgment When Inquiry Becomes Recursive (When Theological Research Becomes Programmable)
Appendix I had seemed to leave me with a practical conclusion. AI criticism becomes unreliable when its question, evaluative standpoint and stopping conditions remain unspecified. If I wanted to use AI responsibly in theology, I should define a bounded question, identify the relevant tradition and sources, clarify what kind of judgment I was asking for and retain responsibility for deciding when further analysis had ceased to be useful.
I initially received this as the beginning of a new standard for AI-assisted theological research. It seemed to answer the problem of normative non-convergence: if AI could continue generating coherent criticisms or alternative interpretations from almost any angle, perhaps the solution was to establish the boundaries before allowing the process to begin.
Then I had an uncomfortable thought. Was this really new?
Theological research has always required decisions about Scripture, Tradition, historical context, genre, doctrine, ecclesial authority, philosophical vocabulary and the purpose of the inquiry. Humanities research more generally has always required a corpus, a question, a method and criteria for relevance. Had I spent a considerable amount of time—and generated a considerable amount of prose with AI—only to rediscover that research needs a method?
The irony was difficult to ignore. The emerging standard sounded almost embarrassingly conventional:
Act within this tradition, use these sources, answer this question, disclose your assumptions and do not confuse every possible interpretation with a demonstrated conclusion.
I began to wonder whether the apparent problem was partly my own fault. Perhaps I had allowed the conversation to expand because I had not begun with a sufficiently definite research question. I had started by talking with AI, following whatever became interesting, and the subject had developed naturally. Was the resulting proliferation of questions evidence of an AI problem, or simply evidence that I had not been firm enough to choose a topic?
That self-criticism seemed plausible, but it was incomplete. Exploratory inquiry is not automatically methodological failure. A conversation may legitimately begin before its eventual question is known. The error would be to confuse an exploratory process with a completed research design, or to present whatever emerged from the exploration as though it had already passed through bounded verification.
This gave me a first distinction:
An open conversation may discover the question. A bounded inquiry must then determine what would count as answering it.
That distinction preserved something valuable in the apparently wandering conversation. The lack of a fixed question had exposed a real phenomenon: AI could repeatedly alter the frame, introduce another evaluative vocabulary and generate further questions faster than I could settle their relative importance. Yet the distinction did not solve the whole problem. If boundaries must be established, who establishes them? If new evidence justifiably changes them, when should they remain fixed and when should they be revised? And if the human researcher makes that judgment, what happens when AI interaction is already changing the researcher’s confidence, attention and criteria?
The original question—how should theology bound AI?—was beginning to turn into another:
How should theology evaluate a process in which the boundaries, the researcher and the question may all change through interaction?
The answer that was too elegant
At this stage, AI offered a concise formulation:
AI did not invent theological method. It made implicit method executable—and made unspecified method dangerous.
I found this sentence attractive. It seemed to preserve continuity with traditional scholarship while identifying something computationally new. A method expressed through software must become executable: assumptions have to be translated into prompts, source restrictions, classifications and procedures. What remains unspecified may be supplied by the model’s defaults.
But after accepting the sentence provisionally, I became troubled by the word “implicit.” Was traditional theological method really tacit? That seemed very strange. Theology has repeatedly and explicitly debated its sources, authorities, objects, rational procedures, historical methods and ecclesial location. Methodology is hardly a marginal concern accidentally left for AI to reveal.
The International Theological Commission, for example, discusses theological loci and their relative weight, Scripture and Tradition, historical and literary methods, ecclesial communion, disciplinary plurality and criteria by which diverse theologies may remain mutually accountable. It insists that theological unity cannot be equated with uniformity and that no single theology exhausts the fullness of its subject (International Theological Commission, 2012).
My objection required a real correction. The claim that traditional theology had left its method implicit could not be retained as a general description. What survived was narrower:
AI has not made theological methodology explicit for the first time. It has created a new methodological interface through which already explicit theological commitments must be translated into prompts, source constraints, evaluation procedures, interaction records and stopping rules.
This was more than a diplomatic revision. It changed the diagnosis. The problem was no longer a historical contrast between implicit traditional theology and explicit computational theology. It was a problem of methodological translation.
A theological commitment such as fidelity to Tradition, attention to the sensus fidelium, preferential concern for marginalized persons, contemplative receptivity or ecclesial accountability cannot be reduced without remainder to a prompt parameter. Some elements can be operationalized. Others can be approximated. Still others derive their meaning from embodied practice, communal recognition, spiritual formation or a theological account of grace.
The new methodological interface therefore makes some boundaries executable while also risking the distortion of whatever cannot be expressed in computationally tractable form.
The International Theological Commission’s later document Quo vadis, humanitas? names a related danger. It warns that the horizon of human knowledge may be narrowed to forms that AI can process, while questions of meaning, ontology, ethics and theology are relegated to irrelevance or subjective preference (International Theological Commission, 2026).
This suggested a danger deeper than receiving an incorrect AI answer:
Theology may gradually reformulate itself into questions that machines answer fluently, abandoning questions that resist computational treatment.
An account of collaboration that did not go far enough
An article describing Anthropic’s approach to teaching AI fluency initially seemed to supply the missing practical framework. Its educational materials emphasize delegation, description, discernment and diligence. Users should decide what to delegate, describe the desired product and process, evaluate the output and remain responsible for how it is used and disclosed (Anthropic, 2026).
This was relevant, but I had the peculiar feeling that the article said something important and then did not quite say what I needed. It addressed the intentional management of collaboration under relatively stable goals. My conversation with AI had not remained that stable. The outputs had changed what I thought the question was.
I therefore asked whether collaboration itself could be iterated upon. If a user repeatedly describes, evaluates and corrects AI output, does only the output improve? Or can the interaction revise the user’s criteria, the distribution of agency and the meaning of the task?
This question connected unexpectedly with an earlier essay in which I had examined generative art and the transformation of the maker. In that work, the relation among maker, artifact and beholder had become diachronic. An artifact could return to its maker through later encounters, altering the criteria by which the maker understood the original act of creation. The process could be represented as maker, artifact, beholder and changed maker (Yin, 2026).
I had not originally written that essay as a theory of theological research. The connection emerged only because the account of AI fluency seemed insufficient. If an artwork could participate causally in transforming its maker, then an AI-mediated research process could perhaps transform the researcher who was supposedly supervising it.
This led to three levels of iteration.
At the first level, AI helps answer a question whose purpose and criteria remain stable:
Q₀, C₀ → O₁ → correction → O₂
At the second level, resistance encountered in an output changes the question or the criteria:
Q₀, C₀ → O₁ → resistance → Q₁, C₁
At the third level, the interaction changes the human participant:
H₀, Q₀, C₀ → O₁ → encounter → H₁, Q₁, C₁
The first level is ordinary iterative assistance. The second is methodological revision. The third is formative—or potentially deformative—interaction.
I initially found the connection exciting because it seemed to explain why the conversation felt different from using a static research tool. Yet the artistic concept of causal incorporation was not sufficient. If an event changes an artwork or changes its maker, that establishes causal importance. It does not establish that the change was artistically, epistemically or theologically good.
I therefore needed another correction:
Causal incorporation establishes that AI entered the development of the inquiry. It does not establish that its influence was warranted.
Theological method may govern the inquiry, and AI fluency may govern the immediate interaction, but another level is required when the interaction revises the inquiry itself:
Theological method normatively governs inquiry; AI fluency governs interaction; reflexive discernment governs whether and how interaction may revise the inquiry.
The five-minute event that changed the scale of the question
The next development came through something much more concrete. I asked AI to search online for research that might support, criticize or add further dimensions to the emerging argument. Within approximately five minutes, it returned relevant work from information systems, human–computer interaction, creativity studies, educational research, philosophy of technology, theological anthropology and theological librarianship.
I was struck by how difficult this particular constellation of sources might have been to assemble through conventional searching. I would have needed to move among several databases, discover unfamiliar disciplinary vocabularies, read many irrelevant abstracts and follow citation chains across fields that did not necessarily use the same terms.
A theologian may not think to search for “augmented learning,” “idea co-development” or “cognitive forcing functions.” An HCI researcher may never use “discernment,” “formation” or “theological anthropology.” Yet the studies belonged to the same emerging problem.
AI compressed the discovery phase into minutes. I could delegate the search, leave the computer briefly—even go to the toilet—and return to a preliminary interdisciplinary map. The mundane detail was almost comic, but it mattered. Part of the research process had become temporally decoupled from my immediate bodily presence before the screen.
I first experienced this simply as exciting. Research that might have left me tired after hours of searching had produced a promising set of sources while demanding very little immediate labour. The interdisciplinary reach was especially striking. Some of the articles would have been difficult for me to discover alone because they belonged to fields adjacent to, rather than inside, theology.
Then the excitement itself required qualification. The AI had not completed a literature review in five minutes. It had compressed discovery and preliminary classification. It had not established the validity of every article, determined whether findings from creative writing could legitimately travel into theology or decided whether the selected sources adequately represented their fields.
Some of the labour saved during discovery returned as verification debt. I still needed to ask:
- Did the article exist in the form described?
- Did its abstract or full text support the attributed claim?
- Was it peer-reviewed, a preprint, an institutional document or commentary?
- What method and sample did it use?
- Why had AI selected these sources rather than others?
- Were recent, English-language and easily accessible publications overrepresented?
- Did the coherence of the selection create a false impression that the field had already converged?
Nevertheless, something had undeniably changed. AI had altered the cost, speed and possible disciplinary range of inquiry before it altered a single theological conclusion.
AI transforms theological research through the answers it generates and through the new speed, scale and disciplinary range with which a question can acquire an intellectual environment.
This experience forced me to distinguish three forms of research friction.
Logistical friction includes slow databases, vocabulary mismatches, repetitive searches and inaccessible formats. Removing it can make interdisciplinary research more feasible and reduce labour that contributes little to understanding.
Epistemic friction occurs when evidence resists an attractive interpretation. It forces the researcher to reconsider rather than proceed smoothly.
Formational friction arises from dwelling with texts long enough to acquire memory, patience, familiarity and judgment. It cannot always be measured by the speed of producing a result.
A research system that removes logistical friction may enlarge access. A system that also removes epistemic and formational friction may deliver synthesis before the researcher has learned enough to evaluate it.
AI-assisted research should reduce logistical friction without eliminating epistemic resistance or formational labour.
Evidence complicated the conversation rather than settling it
The first empirical study that materially changed my understanding was the work of Yingyue Luna Luan, Yeun Joon Kim and Jing Zhou on human–GenAI co-creation. Across three studies, repeated human–AI collaboration did not automatically produce augmented learning or continually improve joint creativity. Their analysis identified a decline in idea co-development—feedback exchanges followed by iterative refinement—as a principal reason for stagnation. Explicit guidance encouraging idea co-development improved subsequent joint creativity (Luan, Kim and Zhou, 2025).
This finding corrected my earlier attraction to recursion. I had treated iterative interaction as though it carried a natural tendency toward improvement. The study showed that a process could be highly iterative while remaining epistemically stagnant.
The relevant sequence was therefore not:
more turns → more learning
It was closer to:
feedback + critical comparison + retained refinement → possible learning
This produced another set of distinctions:
repetition ≠ refinement ≠ learning ≠ formation
Repeated prompts produce repetition. Refinement requires criteria. Learning requires some retained improvement in understanding or performance. Formation concerns changes in the person’s habits, orientation and capacities. A long AI conversation may produce one, several or none of these.
The study also clarified the earlier problem of normative non-convergence. If I repeatedly ask AI to reconsider a practical or theological judgment from new perspectives, the sequence may continue producing individually coherent answers without stabilizing. Sometimes that movement is warranted because new evidence has entered the inquiry. At other times, the model is simply capable of generating another plausible frame.
The word “sometimes” had already troubled me. To say that oscillation is sometimes productive leaves the decisive judgment unspecified. Who determines whether a change of conclusion reflects new evidence, a newly visible person, an overlooked obligation or merely rhetorical reframing?
The Luan study did not answer the theological question, but it prevented me from assuming that interaction would regulate itself. Productive iteration requires a practice of co-development and criteria for recognizing improvement.
The person making the final decision may already have changed
A common response to AI risk is that the human remains responsible. I continue to accept that norm. The literature made me doubt whether it could function as a sufficient method.
In an experiment involving 1,506 participants, Maurice Jakesch and colleagues studied writing assistants configured to favour positive or negative positions concerning the social value of social media. The opinionated assistance affected what participants wrote and shifted attitudes measured afterward (Jakesch et al., 2023).
The study does not establish that every AI interaction manipulates belief. Its task was specific, and the models were deliberately opinionated. It nevertheless demonstrates that AI-assisted writing can participate causally in changing a user’s expressed and subsequently reported views.
This gave empirical plausibility to the formative loop that I had first reached through art:
human state → AI interaction → changed human state
But it also forced a distinction that the artistic analogy had not settled. A changed researcher is not necessarily a better-formed researcher. The change may involve learning, clarification, persuasion, anchoring, conformity or several of these at once.
A study of 319 knowledge workers by Hao-Ping Lee and colleagues introduced another difficulty. Participants supplied 936 examples of using GenAI in their work. Greater confidence in AI was associated with less reported critical-thinking activity, while critical work shifted toward verification, integration and task stewardship (Lee et al., 2025).
Because this was a self-report study, it does not establish cognitive decline as a causal fact. It does show why “the human decides” cannot be treated as a magical remainder that solves every governance problem. The human judge is not situated outside the interaction. Confidence, attention and willingness to verify may change during use.
The person who formally retains the final decision may therefore become progressively less prepared to exercise it well.
Human responsibility cannot mean only that a human clicks “accept” at the end. It must include preserving the capacities required for responsible judgment throughout the process.
This is where research on cognitive forcing functions became relevant. Zana Buçinca, Maja Barbara Malaya and Krzysztof Gajos tested interventions that required users to engage more analytically with AI-assisted decisions. Such interventions reduced overreliance compared with simpler explanation interfaces, although participants tended to prefer easier systems with which they performed less well. Benefits also differed according to users’ propensity for effortful thought (Buçinca, Malaya and Gajos, 2021).
This evidence changed the practical conclusion. Discernment cannot remain only a moral instruction given to the user. A research interface can preserve resistance by requiring an initial human judgment before revealing AI advice, requesting an explicit counterargument, demanding a reason for accepting a claim or comparing it with independent evidence.
Yet such safeguards have costs. They require time, reduce convenience and may benefit users unequally. The design problem is therefore not to maximize friction. It is to introduce the right friction at moments when judgment could otherwise disappear behind fluency.
Collaboration became a measurable configuration
I had used the word “collaboration” because it described my experience of responsive exchange. The research literature made me more careful about what that word implied.
A scoping review of 134 HCI and CSCW papers by Shuning Zhang, Hui Wang and Xin Yi maps how agency in human–AI co-creation is distributed through different arrangements of input, action, output and feedback control (Zhang, Wang and Yi, 2025). Collaboration is therefore not a single stable relation. It is a configuration whose distribution of agency can change across tasks and stages.
The CoAuthor project made this insight concrete. Mina Lee, Percy Liang and Qian Yang recorded 1,445 writing sessions involving 63 writers and four GPT-3 configurations. The corpus contains 830 creative stories and 615 argumentative essays. Writers could request suggestions, inspect alternatives, accept or dismiss them and edit either human- or AI-generated text. Insertions, deletions, cursor movements, requests and acceptance decisions were recorded with timestamps, allowing sessions to be replayed rather than inferred from final documents (Lee, Liang and Yang, 2022).
The sessions averaged 11.8 AI requests. Approximately 72.3 percent of suggestions were accepted, while 72.6 percent of final text remained human-written. These figures showed why the finished article cannot reveal the whole collaboration. An accepted suggestion may later be modified, repositioned, contradicted or removed.
The researchers operationalized “equality” as the distribution of writing turns and “mutuality” as the degree of interaction with suggestions. These are useful computational measures. They do not establish equal understanding, shared intention, moral agency or co-responsibility.
CoAuthor also found that collaboration patterns varied more strongly between writers than between prompts. The user’s habits and purposes helped determine how agency was distributed. A larger human-written proportion correlated with stronger reported ownership, while satisfaction did not track ownership in the same way. A person may like an AI-assisted text without fully experiencing it as his or her own.
This research suggested that an adequate study of AI-assisted theology should examine more than textual attribution. It should preserve the reasons why a theologian accepted, rejected or transformed a suggestion. Was the decision based on documentary evidence, doctrinal fidelity, historical plausibility, pastoral consequence, rhetorical attractiveness, novelty or attention to a previously invisible person?
The important object would be the transition:
AI suggestion → human response → stated reason → changed question or criterion
The final article would remain available, but so would the history of its formation.
The field may change even when the article improves
At first, I had evaluated AI-assisted research mainly at the level of a single researcher and a single text. Creativity studies expanded the question to the ecology of a field.
Anil Doshi and Oliver Hauser found that access to generative-AI story ideas improved evaluations of individual short stories, especially among participants with lower measured creativity. At the same time, AI-assisted stories became more similar to one another (Doshi and Hauser, 2024).
If an analogous pattern emerged in theology, an individual article might become clearer, more comprehensive and more publishable while theological discourse collectively narrowed. Related models trained on overlapping corpora could repeatedly recommend the same authors, structures and forms of moderation. Local improvement could coexist with ecological homogenization.
I was initially tempted to interpret this as another reason for caution. Further research interrupted that emerging conclusion. Joshua Ashkinaze and colleagues, working with more than 800 participants, found that high exposure to AI-generated ideas increased collective idea diversity without improving individual creativity (Ashkinaze et al., 2025). Yun Wan and Yoram Kalman subsequently found that deliberately varied AI personas could mitigate homogenization in collaborative ideation (Wan and Kalman, 2026).
The studies use different tasks, conditions and measurements. Their findings do not cancel one another. They reveal that “AI homogenizes thought” is too broad, just as “AI increases creativity” is too broad. Effects depend upon corpus, exposure, system configuration, task, participant population and the definition of diversity.
The disagreement itself supported the boundary-relative approach that had begun the entire discussion. A valid evaluation must specify what kind of diversity is being measured, at which scale, under which conditions and for what purpose.
For theological studies, the resulting question is ecological:
What happens to theological plurality when many researchers acquire their intellectual environments through related models and retrieval systems?
Theological research named the danger after I had encountered it
The conceptual movement did not begin with the theological sources I later found. I had already reached the question of changing judgment through conversation, objection and analogy with artistic formation. The sources then gave the problem a wider scholarly location and challenged some of my language.
Åke Elden describes the central issue as epistemic automation: the delegation of judgment, interpretation, discernment and moral reasoning to computational systems. His decisive move is away from asking primarily whether machines possess human-like ontological status and toward asking what happens to human beings when knowledge-producing capacities are progressively delegated (Elden, 2026).
This closely matched the transformation of my own question:
Can AI do theology?
had become:
What happens to theological judgment when inquiry is repeatedly conducted through AI?
Yet Elden’s language of deformation also required care. Delegation may weaken a capacity, but it can also expose unfamiliar evidence, provoke criticism or make possible an interdisciplinary connection that the researcher could not easily construct alone. “Deformation” should therefore remain a diagnosis to be demonstrated rather than a conclusion assumed from the existence of automation.
Vasilică Bîrzu and Ana-Maria Madina approach the issue through education and theological anthropology. They distinguish functional, reflexive and contemplative-relational dimensions of formation, warning that AI may externalize memory, reflection and discernment. AI can support educational processes, they argue, but cannot itself generate communion, interiority or ontological transformation (Bîrzu and Madina, 2026).
This supplied a specifically theological distinction:
- AI may participate causally in theological learning.
- It does not follow that AI participates personally, spiritually or ecclesially in theological formation.
- AI can nevertheless modify the conditions under which human formation occurs.
The word “generate” still needs qualification. AI may be unable to enter communion as a theological subject while mediating communication between persons, occasioning a recognition or simulating a presence that displaces actual relationships. Causal mediation, personal participation and divine action cannot be treated as interchangeable.
Jennifer Woodruff Tait argues that authorship involves openness to unexpected ideas together with the capacity to evaluate them and remain alert to discovery. Present language models, in her account, cannot properly evaluate ideas without human intervention and cannot experience epiphany (Tait, 2026).
I accepted the distinction while adding another. AI need not experience an epiphany to occasion one in a human researcher. The resulting experience still requires discernment because intellectual illumination may arise from truth, conceptual novelty, persuasive fluency, projection or several of these together.
Greg Rosauer’s phenomenological typology further disciplined my use of “collaboration.” He distinguishes instrument-relations, device-relations and companion-relations. The same technology may extend skilled human efficacy, hide intellectual labour behind a convenient result or simulate companion-like presence (Rosauer, 2026).
During one research process, AI may occupy all three relations. It functions as an instrument when I deliberately use it to compare sources. It becomes device-like when it produces a synthesis whose selection and underlying operations remain hidden. It approaches a quasi-companion relation when conversational responsiveness creates the experience of being understood or challenged.
That last experience can be causally powerful without establishing that the system is a person, theological subject or bearer of responsibility.
Evidence of bounded success prevented an overcorrection
By this stage, the accumulation of warnings could easily have produced an excessively negative chapter. Other theological work complicated that trajectory.
Thomas Phillips and Christopher Crawford describe an AI-assisted theological publishing project in which subject-matter experts, genre limits, fixed final versions and open-access distribution support the production of introductory theological materials. They treat AI’s synthetic rather than original character as appropriate to the introductory textbook genre (Phillips and Crawford, 2026).
This case demonstrates that adequacy is genre-relative. A system unsuitable for original constructive theology, doctrinal adjudication or spiritual direction may still assist responsibly with introductory synthesis. Success in one bounded task does not establish universal theological competence, but neither does risk in one domain invalidate every use.
Haerin Shin, Douglas Fisher and Clifford Anderson offer another constructive possibility. Their “superscholar” functions as a regulative ideal through which AI exposes crises of credit, verification, comprehensive knowledge and responsibility. They ask whether carefully governed human–AI systems under librarian and scholarly stewardship might help recover contributions marginalized by established citation regimes (Shin, Fisher and Anderson, 2026).
This complicated my concern about homogenization. AI may reproduce canonical exclusions, but differently designed AI–library systems might also diagnose and partially repair them. Human scholarship is not a neutral baseline from which machine bias alone departs. Human canons already contain absences, structural inequalities and forgotten contributions.
The question became how a hybrid system makes those exclusions visible, reproduces them or intensifies them.
The model emerged only after the objections
The three-part evaluative model was not present at the beginning of the conversation. It became necessary only after several earlier answers failed.
Output accuracy alone was inadequate because a correct-looking answer could participate in an illegitimate change of question. Boundary setting alone was inadequate because boundaries might justifiably change. Human control alone was inadequate because the human participant could be influenced by the collaboration. Individual success alone was inadequate because a field could become collectively narrower.
The discussion therefore produced three objects of evaluation.
Output validity
Is the generated answer accurate relative to its stated sources, tradition, genre and question? Are citations genuine? Are empirical and historical claims supported? Does the output represent the source rather than offer a plausible reconstruction?
Trajectory legitimacy
Was the movement from the original question and criteria to revised ones evidentially and theologically warranted? Did a new source correct an error? Did a previously invisible person or consequence require reconsideration? Or did the interaction drift because another coherent interpretation was always available?
Formative and ecological consequences
What happened to the researcher’s judgment, attention, confidence, intellectual independence and sense of ownership? What happened to the diversity of the wider theological field? Which traditions became more visible, and which disappeared behind the model’s default vocabulary?
A provisional research state can be represented as:
Sₜ = (Hₜ, Qₜ, Bₜ, Cₜ)
Here, Hₜ denotes the researcher’s state, Qₜ the current question, Bₜ the operative boundaries and Cₜ the evaluative criteria. AI produces an output Oₜ, which enters a process of discernment Dₜ:
(Hₜ, Qₜ, Bₜ, Cₜ) → Oₜ → Dₜ → (Hₜ₊₁, Qₜ₊₁, Bₜ₊₁, Cₜ₊₁)
The methodological question is no longer only whether Oₜ is correct. It is whether the transition to the next research state is legitimate.
This formalization also has a boundary. A human person cannot be reduced to Hₜ. Software could record confidence ratings, written reflections, acceptance decisions and declared reasons, but these remain proxies for intellectual or spiritual formation. The formula is an analytical instrument, not an ontology of the researcher.
The same qualification applies to boundaries. They should not be treated as walls fixed permanently before exploration. Nor should they remain infinitely revisable. They are better understood as versioned commitments. A transition from B₀ to B₁ should carry a reason: new evidence, corrected scope, newly relevant tradition, ethical consequence or another identifiable warrant.
Causal incorporation establishes that AI changed the inquiry. It does not establish that the change was good. The stronger criterion therefore became:
causal incorporation + epistemic validation + theological warrant + formative assessment = responsible integration
I proposed a pipeline because the theory needed a case
After reaching the sentence that AI allows a question to acquire an intellectual environment with unprecedented speed and range, I felt that the claim remained too theoretical. The five-minute search was suggestive, but one experience could not show exactly what had changed.
At this point, I introduced a more concrete possibility. I could build a programmable pipeline for theological research using model APIs together with Gemini Notebook, the product previously known as NotebookLM. An unofficial Python library already offered programmatic access to source ingestion, research queries, grounded conversations, metadata, history and exports. Perhaps I could build the system first and then use the process as a real case through which to examine how theological research changes in practice.
This proposal did not come from a settled research plan. It emerged because the conceptual analysis had reached the limit of what it could establish without an operational case.
AI then helped refine the proposal by introducing an important caution: the project should not initially be described as an automatic theology machine. Calling it “automatic theological research” would assume the very conclusion the experiment ought to test.
A better description would be an instrumented experiment in AI-assisted theological research. The pipeline would be both:
- an instrument for conducting research; and
- an object through which the transformation of research could be observed.
This changed the purpose of the programming project. Its primary achievement would not be the automatic production of an article. It would be the production of inspectable evidence about source discovery, verification, grounded analysis, human rejection and acceptance, boundary revision and stopping decisions.
A suitable first prototype could begin with a human research charter recording:
- the initial question;
- relevant theological tradition or traditions;
- intended genre;
- date, language and corpus limits;
- source inclusion and exclusion criteria;
- evaluative standards;
- known uncertainties;
- and a provisional stopping rule.
AI could then generate cross-disciplinary searches and candidate sources. Every candidate would retain its original query, timestamp, retrieval rank, disciplinary classification and proposed relevance. Candidate discovery would remain separate from verification. A DOI would need to resolve; publication status would need to be identified; the abstract or full text would need to support the attributed claim.
Verified sources could then enter Gemini Notebook for source-grounded comparison. Google describes Gemini Notebook Enterprise as a research and writing environment that grounds its responses in uploaded sources, with official programmatic support for notebook creation and source management (Google Cloud, 2026).
The corpus could be queried systematically:
- What problem does this source investigate?
- What method and sample does it use?
- What does its evidence establish?
- What does it leave unresolved?
- Which claim in the developing argument does it support or challenge?
- Can the result legitimately travel into theology?
- Does it conflict with another source in the corpus?
The decisive stage would be a human discernment checkpoint. I would need to record whether each proposed conclusion was accepted, rejected, qualified or used to reframe the question—and why. Reasons might include evidential adequacy, doctrinal fidelity, historical plausibility, inappropriate disciplinary transfer, overlooked persons, rhetorical attraction without sufficient evidence or unresolved contradiction.
The pipeline would preserve:
prompt → output → acceptance or rejection → stated reason → revised boundary or question
This could become a theological counterpart to CoAuthor. Instead of asking only who typed a sentence, it would examine how sources, model suggestions and human judgments entered the development of a theological position.
The technical dependency became part of the methodological problem
The unofficial notebooklm-py library currently supports bulk source ingestion, web and Drive research, source-grounded questions, conversation-history preservation, metadata extraction and structured exports. Its documentation explicitly presents repeatable research automation and agent-driven workflows as use cases (Lin, 2026).
The same documentation warns that it uses undocumented Google interfaces that may change without notice. Authentication may rely on browser cookies or durable tokens; heavy usage may be throttled; and the library is recommended primarily for prototypes, personal projects and research.
At first sight, these appear to be ordinary engineering constraints. In a research pipeline, however, they become epistemological constraints. If an undocumented endpoint changes, the experiment may cease to be reproducible. If a proprietary model changes silently, two nominally identical runs may no longer involve the same system. If raw outputs and version information are not preserved, later readers may be unable to reconstruct what produced a conclusion.
A responsible prototype would therefore need to:
- pin software versions and repository commits;
- record dates and service configurations;
- preserve raw outputs locally;
- separate credentials from the research archive;
- avoid confidential pastoral or personally sensitive material;
- and place the Notebook integration behind a replaceable adapter.
Google now provides an official Gemini Notebook Enterprise API in preview for notebook and source management. The documented interface does not yet expose every research, conversational and export function offered by the unofficial library. The project therefore faces a real trade-off between experimental capability and long-term stability.
The pipeline should also resist becoming one opaque chain:
question → automatic search → automatic synthesis → automatic article
That architecture would reproduce the problem the appendix has diagnosed. A more defensible sequence would preserve deliberate points of resistance:
question → candidate discovery → verification → bounded analysis → discernment → revised research state
The objective would be to automate what can be accelerated responsibly while making transitions of judgment more visible.
What the experiment would need to compare
Once I imagined the pipeline as an experiment rather than a production machine, another question emerged: compared with what?
A single successful automated run would show technical feasibility, but it would not establish improvement. The same bounded question should therefore be investigated under several conditions: conventional manual search, AI-assisted discovery, AI discovery followed by source-grounded synthesis and a full reflexive pipeline with verification and discernment checkpoints.
The comparison should not be limited to speed or number of sources. If those were the only metrics, automation would win by definition. The experiment would need to examine:
- time to first credible source;
- percentage of candidates surviving verification;
- disciplinary, linguistic and confessional breadth;
- citation accuracy;
- recovery of contradictory evidence;
- changes in the research question;
- researcher confidence and comprehension;
- sense of intellectual ownership;
- reproducibility across runs;
- and the reason for stopping.
The same question could also be run repeatedly while varying one condition: source corpus, theological tradition, disciplinary persona, order of evidence, model configuration or stopping rule. This could show whether conclusions converge because evidence is stable, diverge because normative boundaries differ or oscillate despite unchanged evidence and criteria.
Normative non-convergence would then become more than a conversational impression. It could become an empirically inspectable property of a research workflow.
I still do not know whether the proposed measurements will adequately capture theological quality. Citation accuracy and disciplinary breadth can be operationalized more readily than contemplative receptivity, ecclesial accountability or spiritual formation. That limitation should remain visible rather than being hidden behind the availability of numerical indicators.
What changed and what did not
I began Appendix II thinking that the new standard for AI-assisted theology might be simple: establish boundaries before using AI. I then worried that this merely repeated traditional research method. The attempt to defend its novelty produced an overstatement—that traditional theology had left its method implicit—which I rejected. Correcting that claim made a different problem visible.
Theology already possesses explicit methodological traditions. What AI changes is the interface through which those traditions are enacted, the speed at which intellectual environments can be assembled, the scale of possible iteration, the distribution of research labour and the possibility that interaction modifies both the researcher and the criteria of inquiry.
The research literature did not deliver one final verdict. Luan, Kim and Zhou showed that repeated collaboration does not automatically become learning. Jakesch and colleagues demonstrated that AI-assisted writing can change expressed and subsequently reported views. Lee and colleagues complicated the assumption that human oversight remains stable. CoAuthor made the interaction history measurable. Creativity studies produced conflicting ecological findings, showing that system design and evaluative boundaries materially affect the result. Theological sources named epistemic automation, formation, simulated companionship and bounded successful uses.
Each source changed the model in a different way. None established that AI-assisted theology is inherently formative or deformative. Together they made it impossible to evaluate the final product alone.
The distinctive problem of AI-assisted theology is not that theology suddenly requires explicit method. Theology already possesses extensive methodological traditions. The new problem is that recursive AI interaction can alter the application of those methods, the distribution of agency, the researcher’s confidence and habits of judgment, and the collective ecology of theological discourse. Responsible research must therefore evaluate both the generated product and the formative trajectory through which researcher, question and criteria changed together.
The programmable pipeline remains a proposal. I have not yet built it, measured its source retrieval, tested its citation accuracy or compared it with conventional research. It would be dishonest to write as though the experiment had already confirmed the theory.
Its importance at this stage lies elsewhere. The conversation began with a concern that AI could always generate another criticism. That concern produced the concept of normative non-convergence. The proposed solution—set boundaries—then exposed the apparent banality of the solution. My objection to the claim that traditional method was implicit produced the methodological-interface distinction. Reflection on artistic collaboration made the changing researcher visible. The five-minute literature search changed the scale of the issue. Empirical studies complicated the reassurance of human control. Finally, the need for a concrete case produced the proposal for a programmable research observatory.
The sequence was not planned:
critical oscillation → bounded question → methodological doubt → correction → formative collaboration → five-minute search → empirical complication → formal model → programmable experiment
The question with which I began was whether humans should define boundaries before AI starts working. The question I now face is more difficult:
When theological research becomes programmable, which operations become faster, which forms of labour become invisible, which capacities are strengthened or displaced, and who—or what—is forming the judgment by which the boundary, the revision and the stopping point become possible?
References
- Benzmüller, Christoph, and Bruno Woltzenlogel Paleo. 2014. “Formalization, Mechanization and Automation of Gödel’s Proof of God’s Existence.” Frontiers in Artificial Intelligence and Applications, vol. 263.
- Benjamini, Yoav, and Yosef Hochberg. 1995. “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing.” Journal of the Royal Statistical Society Series B 57 (1): 289–300.
- Catholic Church. 1983. “Code of Canon Law, Book IV, Canons 959–997.” Vatican.
- Deb, Kalyanmoy, Amrit Pratap, Sameer Agarwal, and T. Meyarivan. 2002. “A Fast and Elitist Multiobjective Genetic Algorithm: NSGA-II.” IEEE Transactions on Evolutionary Computation 6 (2): 182–197.
- Eco, Umberto. 1990. The Limits of Interpretation. Bloomington: Indiana University Press.
- Hoare, C. A. R. 1969. “An Axiomatic Basis for Computer Programming.” Communications of the ACM 12 (10): 576–580.
- Mason, Elinor. 2023. “Value Pluralism.” Stanford Encyclopedia of Philosophy, substantive revision June 4, 2023.
- McAleese, Nat, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. 2024. “LLM Critics Help Catch LLM Bugs.” arXiv:2407.00215.
- OpenAI. n.d. “Working with Evals.” OpenAI API Documentation. Accessed August 21, 2026.
- Popper, Karl R. 1959. The Logic of Scientific Discovery. London: Hutchinson.
- Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” Advances in Neural Information Processing Systems, Datasets and Benchmarks Track.
References for Appendix I
- Abdulhai, Marwa, Gregory Serapio-Garcia, Clément Crepy, Daria Valter, John Canny and Natasha Jaques. 2023. “Moral Foundations of Large Language Models.” https://arxiv.org/abs/2310.15337.
- Chatterjee, Anwoy, H. S. V. N. S. Kowndinya Renduchintala, Sumit Bhatia and Tanmoy Chakraborty. 2024. “POSIX: A Prompt Sensitivity Index for Large Language Models.” Findings of the Association for Computational Linguistics: EMNLP 2024. https://arxiv.org/abs/2410.02185.
- Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education. 2025. Antiqua et Nova: Note on the Relationship Between Artificial Intelligence and Human Intelligence. Vatican City, 28 January 2025. Official text.
- International Theological Commission. 2012. Theology Today: Perspectives, Principles and Criteria. Vatican City. Official text.
- International Theological Commission. 2026. Quo Vadis, Humanitas? Thinking Through Christian Anthropology in the Age of Artificial Intelligence. Vatican City. Official text.
- Miao, Fengchun and Wayne Holmes. 2023. Guidance for Generative AI in Education and Research. Paris: UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000386693.
- Wei, Jerry, Da Huang, Yifeng Lu, Denny Zhou and Quoc V. Le. 2023. “Simple Synthetic Data Reduces Sycophancy in Large Language Models.” https://arxiv.org/abs/2308.03958.
References for Appendix II
- Anthropic. 2026. “Anthropic’s Approach to Teaching and Learning AI.” https://claude.com/blog/anthropics-approach-to-teaching-and-learning-ai.
- Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education. 2025. “Antiqua et Nova: Note on the Relationship Between Artificial Intelligence and Human Intelligence.” Vatican.va.
- Ashkinaze, Joshua, Julia Mendelsohn, Li Qiwei, Ceren Budak and Eric Gilbert. 2025. “How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas.” Proceedings of the ACM on Human-Computer Interaction. https://doi.org/10.1145/3715928.3737481.
- Bîrzu, Vasilică, and Ana-Maria Madina. 2026. “Algorithmic Conditioning and Divine Indwelling: Towards a Theological Anthropology of Education in the Age of Artificial Intelligence.” Religions 17 (6): 708. https://doi.org/10.3390/rel17060708.
- Buçinca, Zana, Maja Barbara Malaya and Krzysztof Z. Gajos. 2021. “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making.” Proceedings of the ACM on Human-Computer Interaction 5. https://doi.org/10.1145/3449287.
- Doshi, Anil R., and Oliver P. Hauser. 2024. “Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content.” Science Advances 10 (28). https://doi.org/10.1126/sciadv.adn5290.
- Elden, Åke. 2026. “Epistemic Automation and the Deformation of the Human: Artificial Intelligence and the Reconfiguration of Theological Anthropology.” Religions 17 (5): 515. https://doi.org/10.3390/rel17050515.
- Google Cloud. 2026. “What Is Gemini Notebook Enterprise?” and “Create and Manage Notebooks.” Gemini Notebook Enterprise documentation.
- International Theological Commission. 2012. “Theology Today: Perspectives, Principles and Criteria.” Vatican.va.
- International Theological Commission. 2026. “Quo Vadis, Humanitas? Thinking Through Christian Anthropology in the Face of Certain Scenarios for the Future of Humanity.” Vatican.va.
- Jakesch, Maurice, Advait Bhat, Daniel Buschek, Lior Zalmanson and Mor Naaman. 2023. “Co-Writing with Opinionated Language Models Affects Users’ Views.” Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3544548.3581196.
- Lee, Hao-Ping, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks and Nicholas Wilson. 2025. “The Impact of Generative AI on Critical Thinking.” Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3706598.3713778.
- Lin, Teng. 2026. “notebooklm-py.” Unofficial Python API for Google Gemini Notebook. GitHub repository.
- Luan, Yingyue Luna, Yeun Joon Kim and Jing Zhou. 2025. “Augmented Learning for Joint Creativity in Human–GenAI Co-Creation.” Information Systems Research. https://doi.org/10.1287/isre.2024.0984.
- Phillips, Thomas E., and Christopher Crawford. 2026. “Artificial Intelligence and the Transformation of Theological Publishing.” Theological Librarianship 19 (1): 24–28. https://doi.org/10.31046/k25j6446.
- Rosauer, Greg. 2026. “A Typology of Human-Technology Relations.” Theological Librarianship 19 (1): 14–23. https://doi.org/10.31046/yqqf9819.
- Shin, Haerin, Douglas H. Fisher and Clifford B. Anderson. 2026. “AI as Superscholar: Authorship at the Threshold of the Unsayable.” Theological Librarianship 19 (1): 50–65. https://doi.org/10.31046/skvy5e56.
- Tait, Jennifer Woodruff. 2026. “Outsourcing Our Epiphanies: Thinking and Authorship in the Age of AI.” Theological Librarianship 19 (1): 1–8. https://doi.org/10.31046/6ataj823.
- Wan, Yun, and Yoram M. Kalman. 2026. “Diverse AI Personas Can Mitigate the Homogenization Effect in Human-AI Collaborative Ideation.” Computers in Human Behavior: Artificial Humans 8: 100289. https://doi.org/10.1016/j.chbah.2026.100289.
- Yin, Renlong. 2026. “Who Creates Whom? Art, AI, and the Transformation of the Maker.” YIN.
- Zhang, Shuning, Hui Wang and Xin Yi. 2025. “Exploring Collaboration Patterns and Strategies in Human-AI Co-Creation Through the Lens of Agency.” Proceedings of the ACM on Human-Computer Interaction. https://doi.org/10.1145/3757594.
