What Counts as Progress in AI-Assisted Theology

The most useful turn in this inquiry came when I rejected a question that seemed to be bringing the discussion into focus. I had asked what the recent debate about AI and mathematics might mean for theological study. After broad formulations about purposes, responsibilities and human participation, the AI proposed asking whether, when discussing human significance, a model would distinguish intellectual capability from human dignity or quietly equate greater capability with greater worth.

The distinction was recognizable, and the question was clearer than asking generally about AI and human significance. Yet it seemed to miss the problem I was trying to reach. My working judgment was that capable AI systems could often understand this distinction and reason within an established theological framework. I wanted to ask what would become urgent under those circumstances. What changes when AI produces theology that is accurate, sophisticated and responsive to objections?

That objection redirected the inquiry. We had been moving toward a question about a possible failure of theological interpretation. I wanted an inquiry that would remain consequential when interpretation succeeded. In developing that objection here, I turn to the relationship between a valuable argument, the person who adopts it, and the institutions that recognize its contribution. These are working questions, rather than conclusions I had already reached in the conversation.

What the mathematical debate made visible

The immediate background was a collection of discussions about AI-assisted mathematics. Reports of major results, enormous computational resources and disputes over research credit had prompted me to ask how other disciplines might change. The report I brought into the conversation described approximately 10,000 concurrent agents working toward a claimed Navier–Stokes resolution in about 88 hours. OpenAI’s own announcement supplies an account of that process and links to a mathematical paper and Lean material. I cite it here as the company’s account; independent validation of the proof remains a separate question. What initially entered my inquiry was the scale of the reported research process (OpenAI, 2026).

At that point, I asked a fairly direct question: could theological research make meaningful use of computational resources on anything like this scale? I was unsure what thousands of agents would be asked to accomplish, how quickly their work could be evaluated, or what would count as a valuable outcome. The issue initially looked like a difference in resources and methods. The subsequent discussions made me consider something less easily measured: a discipline’s capacity to absorb what its tools produce.

The declaration A Severe Misalignment of AI in Mathematics expressed concern about the relationship between solving problems, developing understanding and sustaining the community through which mathematical ideas are transmitted. It also recognized AI’s potential to improve mathematical study. Reading it alongside the surrounding discussions made it difficult to reduce the disagreement to a contest between enthusiastic innovators and frightened traditionalists (Math and AI, 2026).

One account on Zhihu, written under the name 少十七, made that difficulty concrete. The author described the excitement of an AI-generated suggestion that helped a research group approach a problem differently. Later, the same account expressed concern about papers becoming instruments for claiming credit even when their authors had not fully understood them. I read this as first-person testimony, not as a verified description of the entire mathematical community. Its further speculation about corporate motives required evidence beyond the experience reported (少十七, 2026).

What mattered was the change within the account. The author had welcomed AI because it helped the researchers understand something. The later objection concerned a situation in which producing publishable material could cease to involve that understanding. This gave me a reason to question the simple choice between welcoming AI and defending human research. The same person could welcome a discovery and object to the incentives surrounding its production.

There was also a technical distinction that several comments blurred. A formal proof is much more than a machine returning “true.” Proof assistants can check detailed derivations against specified foundations, and proof objects can be checked independently. Knowing how a system discovered a proof is different from checking the resulting derivation. Conceptual explanation raises a further question: what does the proof allow a reader to understand or do? These distinctions prevented me from treating every difficult machine-produced argument as an unexplained answer (Lean Project, n.d.).

Source checking for this article added a useful complication from Professor Timothy Gowers’s discussion of the earlier Leiden Declaration. He considered a future in which AI could produce correct proofs and explain them well, questioned whether mathematicians’ enjoyment could justify preserving existing ownership arrangements, and expressed concern about the possible loss of shared mathematical culture. He also refused to assume that exposition would remain a permanently human advantage. His position helps distinguish taking technical success seriously from accepting every change to the practices surrounding it (Gowers, 2026).

I did not yet have an adequate theological question. The mathematical debate gave me a difficulty to bring back to theology: a field could acquire excellent results while the relationship between those results and its members’ understanding changed. That was more demanding than asking whether AI would become better at completing particular tasks.

Why I initially expected theology to be different

There had been an earlier reason for my concern about human participation. In a conversation with a friend, he suggested that activity benefiting humanity alone lacked ultimate significance unless it participated in a larger cosmic evolution. He also acknowledged uncertainty about what that evolution would mean. I wondered whether this made persons an intermediate stage in a process whose value lay elsewhere. That was my interpretation of the tension, rather than an established account of his position. It led me to ask whether theology offered a different understanding of why human beings matter.

I was inclined to think that theology’s embodied and relational dimensions might make its encounter with AI different from mathematics. Theology concerns persons, communities, practices and forms of life; much theological education also involves learning how to listen, interpret and respond. Mathematics seemed more directly exposed because its objects are abstract and some of its achievements can be expressed in formally checkable terms.

An AI formulation gave this intuition a clearer shape: “AI could profoundly reshape theology without visibly replacing theologians.” I responded that this seemed likely, and then asked whether mathematics might experience both profound change and visible replacement. This was my first attempt to distinguish the two fields’ possible futures. It remained a conjecture. The continued presence of theologians would tell us little, by itself, about how their work had changed.

The mathematical discussions complicated the contrast further. Abstract objects do not make mathematical practice disembodied. The comments about students, seminars and the transmission of methods described a social practice with its own forms of intellectual formation. Conversely, theological research includes substantial textual, historical and argumentative work that can be assisted or automated without first resolving the meaning of religious experience. My emphasis on embodiment still identified something worth examining, but it did not establish theology’s immunity to disruption. I had to look more closely at particular activities.

Professor Harris’s account of his encounter with AI gave the discussion a theological point of entry. In the transcript of his OCTAI Network Voices interview, around 3:24–4:26, he describes expecting to approach AI from a relatively detached theological position and finding that theology entered his questions at every level. His discussion turns toward the responsibilities involved in creating technologies capable of changing human life and the wider world (Professor Harris, n.d.).

My immediate response was to ask whether this showed how fundamentally AI could influence theology. Looking more carefully at the direction of the claim, however, I needed to distinguish two things. Professor Harris was describing theological formation shaping his approach to AI. I was asking whether engagement with AI would, in turn, reshape theology. The second possibility does not follow automatically from the first. His observation made that reciprocal question available to me; it did not answer it.

What survived this correction was the importance of where theological judgment enters. If it is already involved in asking what a technology is for, its role extends into the selection of problems and standards. This helped me move beyond predicting which profession might lose its jobs. Even where the people remain, the questions they pursue and the achievements they recognize may change.

I then wondered whether the questions were still unclear because theology had not adequately defined its own purposes or boundaries. That suspicion was too sweeping. Within Catholic theology, for example, Theology Today offers developed accounts of theological sources, rationality, communal responsibility and the critical use of other disciplines. It relates the plurality of theological methods to the pursuit of knowledge of God. This account does not bind every academic or religious tradition, but it shows that the field is not beginning without methodological resources (International Theological Commission, 2012).

The difficulty was becoming more specific. An institution can articulate a purpose clearly and still lack an adequate way of judging whether a new practice serves it. Saying that education should cultivate judgment does not tell us whether a particular AI-assisted research process develops that judgment. Saying that theology should seek truth does not establish that greater publication volume represents progress toward it.

A question can be precise and still miss the problem

The AI offered a formulation that brought several concerns together:

If AI can supply increasingly excellent intellectual products, what should we deliberately preserve, change, and fund so that people can still develop understanding and meaningfully participate in the practices those products serve?

I brought this question back to theology, but still found the discussion too vague. Preserving understanding and participation sounded reasonable. I could not yet see what, specifically, a theological researcher should investigate. Agreement with the aspiration was leaving the research problem underdetermined.

The formulation also connected with my earlier writing. In my work on boundary-relative computational theology, I had argued that choosing a corpus, representing concepts and assigning authority already involves theological decisions. In a later essay, I had examined how AI might help form the criteria by which its own influence is evaluated. Repeating those arguments would not resolve my dissatisfaction. I needed to discover what the current discussion added to them (Yin, 2026a); (Yin, 2026b).

I therefore asked for greater precision. One formulation offered in the exchange was:

When AI discusses human significance, does it distinguish intellectual capability from human dignity, or silently treat greater capability as greater worth?

This identified a particular interpretive failure that could be investigated. It also made an ethical concern more concrete. But its focus was still on whether a model would mishandle a distinction that theological traditions already provide resources to explain.

My second objection concerned its priority. I was not disputing the importance of distinguishing dignity from capability. I was questioning whether a model’s failure to express that distinction was the central problem before us. A capable model might correctly state a tradition’s account, identify its premises and discuss its implications. That possibility was exactly why I had asked what AI would do to theological research. I wanted to understand what became difficult after granting the competence, rather than repeatedly returning to cases in which it was absent.

This was a working assumption, not a claim that every current model reliably understands every theological tradition. Accuracy, fabricated references and distortions still need examination. Nevertheless, an account of AI’s significance that depends entirely on its mistakes becomes less useful as those mistakes diminish. I wanted to remove an easy source of reassurance and see which questions remained.

The two objections therefore did different work. Asking for precision challenged the breadth of the discussion. Rejecting the dignity example changed the condition under which the next question would have to matter. A narrow question could still miss my concern. Only after this second objection could I state the requirement more clearly: grant that the AI’s theological work is good, and then ask what follows.

An excellent argument and the person who adopts it

To develop that correction, I now find it useful to distinguish producing valuable theological scholarship from possessing theological expertise. What must a researcher understand and be able to justify before adopting an AI-generated argument as their own scholarly judgment? This is a candidate question developed here with AI assistance. Its importance does not depend on the argument containing an obvious mistake.

Consider a hypothetical chapter whose sources are accurately represented, whose argument is substantial and whose responses to objections are convincing. Suppose extensive AI assistance helped produce it. The chapter’s quality should be assessed on its merits. A successful argument does not lose its value simply because its production involved a machine. At the same time, the text alone may provide limited evidence about what its named author can explain, assess or responsibly endorse.

Three judgments therefore need to be distinguished. One concerns the quality of the work. Another concerns the researcher’s contribution to its production. A third concerns what the researcher has come to understand. A person might learn deeply from an argument they did little to originate. Someone else might contribute substantially to organizing a project while possessing only partial expertise in its conclusions. Another might publish sophisticated material that they have barely examined. These possibilities make attribution and assessment more demanding; they do not supply an automatic verdict on AI-assisted work.

The chapter could also be the record of accelerated learning. An AI system might explain difficult sources, expose weaknesses in a draft, introduce relevant objections and help a researcher revise their understanding. The person might become more capable through precisely the assistance that reduced the amount of unaided writing. It would be a mistake to identify learning with the preservation of every previous difficulty. Time spent struggling is not, by itself, evidence of educational value.

Here the earlier objection needs to be applied again. If I reserve explanation, criticism or question formation for humans simply because AI has taken over drafting, I have only moved the assumed boundary. AI may assist with those activities too. The educational question concerns whether the researcher learns through the interaction. A tool’s ability to supply an explanation does not establish that its user understands it; neither does that ability prevent understanding from developing.

The relevant evidence would concern what the researcher can now do. Can they explain the central inference and why its premises matter? Can they assess a serious objection that was not already included in the chapter? Can they recognize when a newly introduced source changes the argument? Can they distinguish what they have established from what remains disputed? These are candidate indicators of judgment, not a complete theory of theological expertise. They would need to be adapted to historical, systematic, practical and other forms of research.

Nor should “personally understand” mean reproducing an entire intellectual process without assistance. Scholarship depends on translations, editions, testimony and the work of specialists. A demand for total self-sufficiency would misdescribe the activity we were trying to protect. The more defensible requirement concerns informed dependence: being able to identify where one relies on others, why that reliance is warranted, and what would require reconsidering it.

Even that requirement leaves difficult cases. A researcher may grasp a conclusion through explanations supplied by the same system that produced it. Further explanation can genuinely deepen understanding, but agreement within one continuing interaction is not an independent check. Returning to primary sources, discussing the argument with other scholars and examining objections from outside the initial exchange can broaden the basis of judgment. None guarantees correctness. Together, they make reliance more open to challenge.

What would count as an advance in theology?

A related question concerns research itself. If competent theological argument becomes easier to produce, institutions will need to identify more clearly what a publication contributes. A defensible interpretation can be worthwhile, but the ability to generate another one does not automatically establish its significance. The question becomes concrete when a journal, department or funding body must decide which work deserves sustained attention.

Some contributions can be specified without settling a universal definition of theology. A study may establish something previously unknown about a source, resolve an interpretive difficulty, show why an influential argument fails, or bring evidence into relation that changes an existing conclusion. It may also make neglected scholarship available for examination. These are different achievements. Their value depends on the question and the relevant standards of evidence, rather than on the quantity of sophisticated language produced.

An essay coauthored by Johan Commelin, Mateja Jamnik, Rodrigo Ochigame, Lenny Taelman and Professor Akshay Venkatesh provided a useful parallel. Their formulation that “means reshape ends” concerns the ways technology can change the problems pursued and the forms of mathematical work valued. It also recognizes disagreement within the community about those values. I take its application to theology as an argument to develop, rather than as something their mathematical analysis has already demonstrated (Commelin et al., 2026).

For theology, the danger would be a gradual narrowing of significance around what a system makes convenient to generate, compare and display. Yet convenience could also enable better research. Translation might make a previously inaccessible debate available; a wide search might reveal evidence that unsettles a familiar interpretation. The fact that a method changes an agenda is insufficient to condemn the change. We need reasons for judging whether the new agenda addresses something worth understanding.

This gives my earlier question about defining the question a more precise consequence. Suppose one project asks AI to reconstruct a doctrine faithfully within a tradition, while another asks whether that doctrine needs revision. Fidelity to the existing formulation serves different purposes in the two projects. Treating it as the decisive test in both would prejudge the second inquiry. AI could help formulate either project, but the researcher would still need to justify what was being investigated and why the chosen evidence could answer it.

Making the boundaries explicit would help others examine that judgment. It would not establish their adequacy. A transparent corpus can still be too narrow for the claim made about it, and changing the corpus can be a warranted response to evidence. I would therefore want to examine the reasons for retaining or revising a boundary. The question is sharper when attached to a particular claim than when posed as a general demand to define theology once and for all.

There is also a limit to the analogy with proof checking. Theological research includes claims that can be tested against documents, historical evidence or logical relations. It also includes disputes about the adequacy and authority of the assumptions through which those materials are interpreted. The absence of a single procedure for adjudicating all such disputes does not license weak argument. It makes the specification of the claim, its evidence and its limits especially consequential.

This also changes how I would return to my earlier question about thousands of agents. A larger search for sources, objections and connections could be valuable. Its value would depend on the question being pursued and on how the results were assessed. A theological institution need not reproduce the most expensive mathematical experiment to benefit from AI. Before asking how much computation it could afford, it would need to identify what contribution that computation was expected to make.

What institutions would have to sustain

The relationship between learning and research output gives these questions an institutional urgency. A student’s early project may simultaneously produce a modest result, develop expertise and provide evidence for further opportunities. If AI changes the value or cost of the result, those other functions still need attention. They may require different forms of assessment and support. The transition cannot be evaluated solely by counting papers or estimating time saved.

Antiqua et nova already treats education as the formation of persons and discusses both the benefits of AI support and the risks of dependence that weakens judgment. Its educational account supplies a normative starting point. Whether a particular course or research practice achieves those aims remains an empirical and pedagogical question (Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education, 2025, §§77–84).

A faculty could examine learning through discussion of source choices, responses to unfamiliar material and explanations of how objections changed an argument. Journals could make the claimed contribution and the distribution of intellectual work clearer. Funders could support reliable editions, translations, teaching and shared resources alongside new arguments. These are directions for institutional experimentation. Their effectiveness should be assessed, rather than assumed because they use the language of human formation.

Access also matters. Oxford’s BiblioTech AI Research Initiative explicitly aims to improve access to theological resources and the visibility of Majority World scholarship through translation and search. It illustrates a research agenda in which AI is being developed to address inequalities in whose work can be encountered. The initiative’s objectives should be distinguished from demonstrated outcomes, but they show why increased computational mediation need not mean accepting the existing distribution of scholarly influence (Ian Ramsey Centre, n.d.).

The possibility I would want to investigate is that access to theological material could widen while control over important research tools becomes more concentrated. Those developments could occur together. Better access would be a real benefit; dependence on a small number of platforms would require a separate judgment. Neither the promise of inclusion nor the concern about concentration should be allowed to erase the other.

Theology’s traditions of formation do not guarantee that its institutions will respond well. A faculty might affirm the value of wisdom while rewarding mainly the speed and volume of publication. A system might reproduce that affirmation accurately and still be deployed in ways that make thoughtful engagement harder to sustain. This is one reason the assumption of competent AI changes the inquiry. Correct theological content does not establish that the practice surrounding its production is well ordered.

The question returns to the conversation

There is a risk in the direction we reached. Questions about degrees, expertise and publication could reduce the discussion to the preservation of an academic profession. That would leave too much of theology outside the frame. Within the Christian account I had been considering, inquiry concerns God, creation and salvation, and its truth is not produced by institutional approval. Human participation matters within that concern; it cannot simply replace it.

I therefore cannot judge AI-assisted theology only by whether it makes researchers feel engaged or intellectually fulfilled. A satisfying process can support a mistaken interpretation. Equally, a sound argument can emerge through a process whose educational benefits are limited. The quality of theological claims, the development of the researcher and the sustainability of the community need to be evaluated without assuming that improvement in one guarantees improvement in the others.

The conversation itself provides a small, incomplete instance of this problem. I brought the mathematical discussions, Professor Harris’s remarks and my questions about theology. The AI supplied formulations that I could examine. I objected when they remained too broad, and again when a narrower proposal seemed to identify the wrong priority. The more developed distinctions in this article are a further step in that AI-assisted inquiry, rather than a transcript of conclusions I had already reached.

Asking for a second pass on the article made this issue concrete again. A coherent retrospective account could make it seem that, when I rejected the dignity question, I already had the distinctions developed here in mind. I had not stated them that way. At that point I could identify an inadequate question more clearly than I could supply a better one. The writing helped articulate possibilities that the objection had opened.

Preserving that sequence matters to authorship. It would be inaccurate to present every formulation as something I had independently conceived before the exchange. It would also be inaccurate to describe the process as accepting a completed answer. My interventions changed the task. That is evidence of a contribution, although it does not establish that my objections were correct or that I have mastered the resulting argument. The fluency of the finished article cannot answer those questions on my behalf.

The final questions are therefore working questions: what demonstrates that an AI-assisted researcher has acquired theological judgment, and what demonstrates that an AI-assisted publication advances theological knowledge? Their distinction is clearer than the broad question with which I began. They still require sustained work on particular forms of inquiry, sources and institutional practices. They should not become an excuse to postpone using AI until a complete theory of theology has been agreed.

What changed most was the condition I wanted the inquiry to survive. I no longer wanted its importance to depend on AI misunderstanding a basic theological distinction. I wanted to examine a situation in which the sources were read well, the reasoning was strong and the assistance was genuinely useful. Under those conditions, the questions about what I understand, what I endorse and what I have contributed remain. This article is one of the places where I will have to answer them.

References

Commelin, Johan, Mateja Jamnik, Rodrigo Ochigame, Lenny Taelman, and Akshay Venkatesh. 2026. “Shaping the Future of Mathematics in the Age of AI.” arXiv:2603.24914, version 2, 8 May.

Dicastery for the Doctrine of the Faith and Dicastery for Culture and Education. 2025. Antiqua et nova: Note on the Relationship Between Artificial Intelligence and Human Intelligence. 28 January. Especially §§77–84.

Gowers, Timothy. 2026. “Thoughts about the Leiden Declaration.” Gowers’s Weblog, 26 July.

Professor Mark Harris. n.d. “OCTAI Network Voices: Professor Mark Harris.” Video interview. Passage discussed: approximately 3:24–4:26, using the transcript included in the materials for this inquiry.

Ian Ramsey Centre. n.d. “Faculty of Theology and Religion Launches the BiblioTech AI Research Initiative (BAIRI).” University of Oxford. Accessed 15 September 2026.

International Theological Commission. 2012. Theology Today: Perspectives, Principles and Criteria. Approved in 2011; published in 2012. Especially §§80–81.

Lean Project. n.d. “Introduction.” Theorem Proving in Lean 4. Accessed 15 September 2026.

Math and AI. 2026. “A Severe Misalignment of AI in Mathematics.” Collective declaration. Accessed 15 September 2026.

OpenAI. 2026. “On the Navier–Stokes Millennium Prize Problem.” 8 September; updated 10 September. Company announcement and links to associated research materials.

少十七. 2026. Answer discussing AI-assisted mathematics, research credit and mathematical training. Zhihu, edited 12 September. First-person account reproduced in the discussion materials.

Yin, Renlong. 2026a. “Toward a Boundary-Relative Computational Theology: What Must Theology Become to Be Verifiable?” 21 August.

Yin, Renlong. 2026b. “When the Proxy Writes the Proxy: The Tool, the Worker, and the Unfinished Good.” 3 September.