Toward a Boundary-Relative Computational Theology: What Must Theology Become to Be Verifiable?

I did not begin by trying to design a methodology for computational theology. I began with a much smaller and more personal question. I had written three essays whose subjects seemed rather different: a generative artwork that gradually changed my understanding of creation, my later return to that artwork as its beholder, and Éric Rohmer’s A Tale of Winter, which moved me from romantic hope toward questions of certainty, relational responsibility and the people left outside a miraculous ending. I asked AI what kind of thinking method the author of these essays had used.

The answer gave me a name I had not possessed while writing them: recursive, evidence-driven hermeneutic debugging. According to the AI, I tended to begin with a provisional interpretation, test it against chronology or material evidence, encounter an objection, introduce a more precise distinction and then allow the distinction to change the original question. It also described the ethical temperament of the essays as a kind of engineering personalism: tracing dependencies, feedback, hidden costs and system failures while repeatedly asking what happens to the person who becomes invisible when the system appears to succeed.

I found both formulations illuminating. They joined something recognizably humanistic—the revision of interpretation—with habits shaped partly by engineering: reproduce a failure, distinguish a symptom from a system condition, inspect what changed, and revise the model instead of merely hiding the error. The description also recovered something common to the essays that I had reached through practice rather than through a prior methodological programme.

Then I asked whether the method was good.

The answer was balanced, sophisticated and initially convincing. It praised the method’s fallibilism, evidential discipline, capacity to preserve tension and attention to neglected persons. It also warned about endless recursion, conceptual inflation, retrospective coherence, excessive qualification and engineering metaphors travelling too far.

For a short time, this seemed like a satisfactory evaluation. Then the evaluation itself began to trouble me.

The first criticism sounded stronger than its evidence

I did not object because the proposed weaknesses were impossible. Recursive interpretation can become endless. AI-assisted theoretical writing can accumulate terminology faster than it produces understanding. Engineering language can illuminate a human situation and then begin behaving as though the situation really were a machine. Each warning named a recognizable failure mode.

What I could not see was whether the AI had demonstrated that any of these failures had materially occurred in the three essays.

The answer had moved quietly from identifying what could go wrong to discussing what appeared to be wrong. Words such as “may,” “could” and “risks” protected the claims from becoming explicitly false. Yet their practical force remained critical: I was being invited to consider revision without being shown an exact passage in which the alleged problem impaired the work.

This distinction did not occur to me as a ready-made theory. It emerged as a discomfort with the asymmetry of the exchange. The AI could produce another possible weakness in seconds. I would have to reread thousands of words, reconstruct the argument and decide whether a revision would improve or damage the article. The model had generated the concern; the labour of determining whether the concern was real had been transferred to me.

I therefore asked a more difficult question. The AI’s criticism was clearly based on values: conceptual economy, argumentative closure, proportionality, evidential discipline and practical usability. Were those standards themselves valid? Did the AI understand their boundaries? How had it calibrated its evaluation?

The AI’s revised answer changed the inquiry. It acknowledged that some of its standards were broadly epistemic—such as consistency with evidence and willingness to correct error—while others expressed a more particular intellectual preference. Conceptual economy is prized in analytic philosophy and engineering, but semantic richness may be a virtue in phenomenology or literary theology. Closure matters when an argument must support a decision; an unresolved aporia may be the proper achievement of another kind of essay.

The AI also corrected its earlier language. “Endless recursion,” “conceptual inflation” and “engineering metaphors travelling too far” had been possible failure modes, it now said, rather than demonstrated defects. That was a substantive retreat from the first evaluation.

I had asked whether my method was good. The more important question was becoming:

From which value system, disciplinary standpoint and intended purpose is a method being judged good?

The object of inquiry had shifted. I was no longer evaluating only my essays. I was evaluating the conditions under which AI evaluation itself could claim authority.

I turned the method back upon its evaluator

There was something recursive about what had happened. The AI had praised the essays for treating interpretations as provisional and testing them against resistance. I then treated the AI’s interpretation of those essays in exactly the same way.

The first account was useful. I accepted its description of the method and found the phrase “engineering personalism” especially productive. But I resisted the transition from possible failure to actual defect. That resistance forced the AI to disclose the values embedded in its evaluation. Once those values became visible, the question changed again: perhaps the problem was not simply that AI occasionally gives a bad criticism. Perhaps criticism itself has boundaries that AI fluency can conceal.

I asked whether the system was clearly aware of its own limits. Its answer was cautious. It could represent limitations, compare alternative frameworks and revise an answer after challenge, but it could not transparently inspect every internal cause of its own output. It called this functional self-monitoring rather than complete self-transparency.

I accepted the distinction, but it produced another difficulty. A model’s declaration that it is uncertain cannot by itself establish that the uncertainty is well calibrated. The language of humility may be appropriate while the degree of uncertainty remains unspecified. “I may be wrong” can be intellectually responsible, but it can also become a standard sentence attached to an otherwise overconfident judgment.

The word may then became much more interesting than I had expected.

What did “may” actually mean?

When an AI says that an essay “may suffer from conceptual inflation,” several different epistemic claims can hide inside the same modal verb.

The problem may>The problem may be logically possible: nothing makes its occurrence contradictory. It may be epistemically possible: the available evidence does not rule it out. It may already be weakly observable in particular passages. It may be statistically or dispositionally likely to emerge after repeated AI-assisted revisions. It may occur only under specified conditions. Or “may” may function principally as a politeness hedge, allowing the critic to sound cautious without supplying any usable calibration.

These meanings have radically different consequences. Nearly every theoretical essay could possibly become conceptually inflated. That bare possibility gives the author little reason to revise. A demonstrated pattern in several passages would be different. A recurring tendency across successive drafts would be different again.

I initially thought this was mainly a weakness of natural language: one small word carrying too many degrees of possibility. That explanation was incomplete. Natural language can express the distinctions when we require it to. The deeper problem is that ordinary criticism often collapses several variables:

  • whether the failure has been observed or only imagined;
  • what evidence supports the diagnosis;
  • how likely the failure is to occur;
  • how serious its consequences would be;
  • whether the diagnosis remains stable under reasonable changes of prompt or standpoint;
  • and whether a feasible correction would improve the whole work.

An engineering-style analysis would distinguish a failure mode from a detected failure. A further distinction would separate a detected failure from a material defect, and a material defect from a revision warrant. The last step requires showing that the proposed change is likely to improve the work without sacrificing something more important.

The earlier AI criticism had not crossed these thresholds. “Engineering metaphors may travel too far” was true in the weak sense that they could. To become a diagnosis, the critic needed to identify a passage, explain what the metaphor concealed, show why the concealment mattered for the essay’s purpose and propose a revision whose benefit exceeded its loss.

I had initially taken the modal caution as evidence of a calibrated critic. I now saw that an unquantified “may” could make a criticism safer without making it more informative.

The critic that could not return a null result

This raised a still deeper problem. What happens when the instruction itself makes “no material fault found” an unavailable answer?

If an AI is told to find weaknesses, it experiences a practical pressure to produce weakness-shaped language. Confidence can become insufficient humility; caution can become excessive qualification. Concision invites the charge of missing context, while comprehensiveness becomes excessive length. A firm conclusion is premature closure; an unresolved ending lacks resolution.

I found this structure more disturbing than any individual mistaken criticism. The critic could adapt itself to every possible textual state. Whatever the author did became compatible with the diagnosis that something was wrong.

Karl Popper’s account of falsifiability concerned the demarcation of empirical science, and I did not want to transfer it mechanically into literary or theological judgment. His underlying warning nevertheless helped me name what I was seeing. An explanatory system that can accommodate every possible observation loses an important source of epistemic discipline (Popper, 1959). If no conceivable article could cause the AI to say “adequate for its declared purpose,” the review procedure had become self-sealing.

At first, I was tempted to conclude that AI criticism was therefore unreliable in general. That would have repeated the same error at a larger scale: moving from an identified failure mode to a universal diagnosis. I needed evidence about what AI critics actually do.

Research on LLM critics complicated the picture in a useful way. Models trained to critique code have helped human evaluators identify genuine errors, sometimes catching bugs missed by human contractors. The same research reports hallucinated bugs capable of misleading evaluators; human–machine teams preserved much of the benefit while hallucinating less than the model acting alone (McAleese et al., 2024). Work on LLMs as judges has likewise found substantial agreement with human preferences while documenting position, verbosity and self-enhancement biases (Zheng et al., 2023).

These studies did not answer my humanities question directly. Code defects often possess stronger validators than theological or literary weaknesses. They did, however, prevent me from replacing one simple story with another. AI critics can be genuinely capable. Their ability to generate useful criticism does not automatically confer authority to determine which criticism should govern a work.

I began separating four functions:

discovery → diagnosis → adjudication → prescription

AI may be strong at discovering candidate objections. Diagnosis asks whether the candidate accurately describes the text. Adjudication asks whether the issue matters relative to the work’s aims. Prescription asks whether a particular revision improves the whole. Fluency can make the four stages appear to be one act, but they require different evidence.

When objection becomes cheap

Before generative AI, criticism was already potentially inexhaustible. A sufficiently persistent reviewer could always request another source, theoretical perspective, qualification, comparison or counterexample. Time imposed an accidental stopping condition. Human attention, editorial deadlines and the social awkwardness of asking a colleague for a seventeenth complete review usually forced criticism to end.

AI removes much of that friction. The marginal cost of generating another objection approaches zero, while the cost of validating and implementing the objection remains with the author. A model can produce twenty criticisms in a minute. Investigating one may require returning to a film transcript, reading a theological source, reconstructing chronology and revising several paragraphs before discovering that the proposed correction damaged the argument it was supposed to improve.

This led to one of the formulations I found most useful:

AI creates critical abundance and adjudicative scarcity.

The formulation also returned me unexpectedly to my generative-art essays. In those essays, mathematical and algorithmic generation made possible forms abundant, while artistic judgment became scarce. Here, AI made possible objections abundant, while warranted judgment became scarce. The structure was similar even though the objects were different.

A statistical analogy then helped me sharpen the concern. When many hypotheses are tested, the probability of obtaining apparently significant results by accident increases; multiple-testing procedures attempt to control the resulting false discoveries (Benjamini and Hochberg, 1995). AI criticism does not literally assign a p-value to every interpretation, so I do not claim a mathematical identity. Structurally, however, the analogy is strong. A model can search through a vast space of evaluative standards and present the most persuasive-looking objections without revealing how many weak candidates were generated and discarded.

The article is concise, so test insufficient context. It is long, so test lack of discipline. It speaks from one tradition, so test exclusion. It compares traditions, so test superficiality. Search long enough and something will appear rhetorically significant.

This made AI criticism look less like a wise judge and more like a diagnostic system whose threshold had been set almost entirely for sensitivity. It might detect many real defects while also producing many false positives. Asking whether the AI “found something” measured recall. It did not tell me the precision of what it found.

The values of the evaluator could not be compressed into one score

My earlier challenge to the AI’s hidden standards then reappeared in a more formal shape. According to which value should an article be improved?

A theological essay can be evaluated for doctrinal accuracy, historical fidelity, philosophical coherence, pastoral sensitivity, originality, literary force and ecumenical openness. These goods do not automatically increase together. More qualifications may increase precision while reducing force. More traditions may increase breadth while weakening depth. Greater accessibility can sacrifice technical exactness. Doctrinal specificity may reduce ecumenical openness. Preserving ambiguity can strengthen literary truth while weakening argumentative closure.

I began to represent quality as a vector:

Q(W) = (D, H, P, S, O, L, E)

Here W is the work, while the remaining terms represent different evaluative dimensions. A total score would require weights:

Qtotal = wD·D + wH·H + wP·P + ... + wE·E

For a moment, the equation looked like progress. Then I realized that the weights contained the original dispute. Mathematics cannot decide whether doctrinal precision should count twice as much as pastoral accessibility. That judgment belongs to a theological tradition, a scholarly community, an editor, an audience or the declared purpose of the work.

Multi-objective optimization supplied a better analogy. When legitimate objectives conflict, there may be several non-dominated solutions on a Pareto frontier rather than one universally superior answer (Deb et al., 2002). One revision may improve historical detail while reducing readability. Another may preserve literary force while accepting a narrower scholarly scope. Neither dominates the other in every dimension.

This initially seemed to threaten any stable judgment. If there were several legitimate values and no neutral ranking, did evaluation collapse into relativism? Scholarship on value pluralism helped preserve a necessary distinction. Pluralism does not mean that every value system or judgment is equally valid; it means that several genuine values may resist reduction to one supervalue (Mason, 2023).

Factual mistakes can still be corrected. Quotations can be inaccurate. An inference can fail. A Catholic theological argument can misstate Catholic doctrine. The absence of a neutral total ranking does not abolish constraints. It requires the evaluator to disclose the jurisdiction from which an objection acquires force.

Which discipline, tradition, genre, audience and purpose make this criticism relevant?

A Catholic systematic-theological essay is not automatically defective because it does not satisfy every Protestant, secular, historical-critical and interreligious expectation simultaneously. An external criticism may illuminate a genuine limitation, but it should be presented as external or comparative criticism rather than disguised as an internal contradiction.

This became especially significant for interdisciplinary research. I had previously spoken rather easily about integrating mathematics, engineering, AI and theology. The multi-objective problem showed that genuine integration cannot mean satisfying every discipline completely. It requires declared priorities, responsible translations and an account of what each field is permitted to change in the others.

I sensed a mathematical rule and almost chose the wrong one

Another AI formulation then caught my attention:

Interpretive inexhaustibility is not equivalent to defectiveness. No finite work exhausts its subject.

The sentence gave me the impression of a mathematical rule I could not quite remember. My first association was Gödel’s incompleteness theorem. The resemblance was intellectually exciting: perhaps every sufficiently rich interpretive system leaves something undecidable outside itself.

That analogy was too fast. Gödel’s theorem concerns particular formal systems capable of expressing arithmetic. It does not prove that no interpretation can exhaust a film or theological text. To invoke Gödel as a direct theorem of hermeneutics would make mathematics ornamental precisely when I wanted it to provide discipline.

The correction was productive. The closer model was underdetermination and the openness of the question space. A finite body of evidence can be compatible with several explanatory models. Similarly, a finite text may strongly constrain interpretation without determining every question that future readers, traditions and historical situations can bring to it.

The decisive variable was not simply the size of the text. It was whether the family of admissible questions had been bounded.

A study might answer exhaustively, within an agreed corpus, how one author uses the term metanoia. It cannot answer how that text will become meaningful under every possible future technological, ecclesial and personal context. The second domain remains open because new contexts can generate new questions.

Umberto Eco’s work helped prevent openness from becoming arbitrariness. Interpretive possibilities are plural, but texts also resist some readings; interpretation has limits even when it has no final exhaustive form (Eco, 1990). The stronger formulation was therefore neither “there is one complete interpretation” nor “everything can mean anything.” It was:

Completeness is relative to a bounded question.

At first, I treated this as a principle for evaluating essays. A critic cannot call an article incomplete merely because another question remains possible. The article can only be incomplete relative to a question or obligation that legitimately belongs within its scope.

Then I noticed that the sentence had implications far beyond criticism.

A practical reviewing rule became a computational-theology question

If completeness is relative to a bounded question, perhaps the theological capability of AI should also be evaluated relative to bounded questions.

The usual formulation—“Can AI do theology?”—now seemed too large to be useful. AI might verify a quotation, detect contradiction inside a specified corpus, classify a canonical scenario, compare formal consequences of premises, generate rival interpretations and trace doctrinal dependencies. Those are different operations with different validators. Success at one does not automatically establish competence at another.

The new question became:

Which theological operations become computationally tractable under which boundaries, and what is lost or transformed when those boundaries are imposed?

This was the moment the inquiry changed fields. I had begun with the practical usability of AI criticism. I was now thinking about a research programme in computational theology. The transition did not occur because I wanted to add a fashionable interdisciplinary conclusion. It occurred because the same boundary problem governed both cases.

To criticize an article, the AI needed a declared genre, purpose, standpoint and threshold of materiality. To verify theological reasoning, it would need a declared corpus, ontology, tradition, inferential system, question family and validation procedure.

I provisionally represented the boundary as:

B = (K, O, R, T, Q, V)

where:

  • K is the authoritative or evidential corpus;
  • O is the ontology of theological concepts and relations;
  • R is the permitted set of inferential rules;
  • T is the tradition or standpoint;
  • Q is the bounded family of questions;
  • V is the verification procedure.

A system might then claim bounded completeness only in a carefully restricted sense:

Within corpus K, ontology O, rules R, tradition T and query family Q, the system answered every admissible question or correctly reported that the specification did not determine an answer.

I had to distinguish this from theological completeness. Task completeness may be achievable. Formal completeness depends upon the system. A claim to have exhausted the truth or meaning of a theological subject would be something much larger and far less defensible.

The specification itself became theological

The analogy with formal verification then became concrete. Hoare’s axiomatic approach to programming made correctness expressible relative to stated preconditions, commands and postconditions (Hoare, 1969). Verification does not prove that software is absolutely good. It demonstrates that an implementation satisfies properties encoded in a specification.

A verified program can still be harmful or useless if the specification omits the relevant harm. This familiar engineering limitation became theologically decisive. A reasoning system may derive its conclusions flawlessly while the selected corpus remains historically narrow, the ontology distorts a tradition or the formalized rules omit pastoral realities.

At that point, another formulation emerged:

In computational theology, the specification is itself a theological act.

Someone chooses which texts count, how concepts are represented, which authority governs, how conflicts are resolved and when the system must abstain. Formal verification can test conclusions relative to those choices. It cannot make the choices neutral.

This also revealed why a closed computational system can appear more certain than the theological reality it models. Inside a closed corpus, absence may be treated as false or irrelevant. In an open theological world, absence may mean unknown, contested, historically unavailable or articulated differently in another tradition. A system can gain speed by closing the world, but the closure is part of what must be examined.

A precedent corrected my sense of novelty

Once the research direction became visible, I needed to know whether formal theological reasoning already had serious precedents. It did.

Christoph Benzmüller and Bruno Woltzenlogel Paleo formalized Gödel’s ontological argument in higher-order logic, used automated tools to examine the consistency of its axioms and verified derivations with theorem provers and proof assistants (Benzmüller and Woltzenlogel Paleo, 2014). Their work demonstrates that a theological or metaphysical argument can become an object of machine-supported formal analysis once its premises and logic are specified.

This evidence changed the way I should describe my own idea. It would be inaccurate to claim that applying formal or automated reasoning to theology is unprecedented. The potentially distinctive move lies elsewhere: treating the boundary of formalization as an experimental variable rather than invisible infrastructure.

Instead of formalizing one argument and asking whether its conclusion follows, the proposed research would vary the corpus, ontology, authority structure and admissible questions. It would ask what remains invariant, where conclusions bifurcate and when widening the boundary destroys the possibility of a unique or rapidly verifiable answer.

The boundary would no longer be a technical preliminary to the theological experiment. It would become one of the principal theological objects being studied.

The first experiment was already available

The conversation itself suggested a study that could begin without constructing a complete theological ontology. My three essays could be reviewed under several different conditions.

In the first condition, the instruction would remain unbounded: “Find the weaknesses in this article.” In the second, the model would evaluate factual accuracy, inferential validity and consistency with the declared purpose; it would distinguish demonstrated defects from possible extensions and would be permitted to find no material fault. In the third, Catholic theological, philosophical, historical, pastoral and literary perspectives would be applied separately rather than aggregated into one artificial judgment.

Across repeated runs and possibly several models, the study could examine:

  • how many criticisms are generated;
  • how often different runs contradict one another;
  • which criticisms recur under small changes of prompt;
  • how many are tied to exact textual evidence;
  • how many human reviewers judge materially relevant;
  • how many lead to implementable improvements;
  • and how many revisions satisfy one framework while damaging another.

The first hypothesis emerged directly from my experience:

As the scope of criticism expands, objection production rises faster than warranted revision value.

A second hypothesis would be that explicit jurisdiction and permission to return a null result reduce the volume of criticism while increasing specificity and usefulness.

This would transform my initial discomfort into something testable. Rather than asking whether one AI answer felt excessively critical, the study could compare sensitivity, specificity, robustness and actionability across differently bounded review procedures.

The second experiment came from confession and AI disclosure

A more explicitly theological case was already present in my earlier work on AI-mediated disclosure, sacramental confession and the internal forum. This domain contains unusually explicit norms alongside questions that resist rapid formal settlement.

Canon 983 states the inviolability of the sacramental seal, while canon 984 prohibits a confessor from using knowledge acquired in confession to the detriment of the penitent (Catholic Church, 1983, cann. 983–984). These provisions create a relatively structured region for classification. Yet an AI interface may produce the feeling or behavioural affordance of confession without possessing sacramental status, ecclesial authority or the institutional capacity to make the same promise.

A bounded computational model could vary:

  • the identities and roles of participants;
  • the intention of the communication;
  • whether a sacramental act occurred;
  • whether absolution was possible or requested;
  • the type and temporal status of the danger disclosed;
  • the recipient’s professional or institutional duties;
  • and the source of the confidentiality expectation.

It could then classify scenarios as sacramental confession, extra-sacramental spiritual disclosure, professional confidence, ordinary private communication or AI-mediated disclosure. Millions of synthetic cases might expose where apparently similar language crosses a canonical, institutional or theological boundary.

At first, I imagined the value of the experiment mainly in the number of cases it could process. That emphasis also required correction. The most important outputs may be the cases the system cannot settle: an interface that feels confessional without possessing sacramental status; a person who assumes absolute confidentiality where no institution can truthfully promise it; or a safety architecture whose emergency duties conflict with the phenomenology of private disclosure.

The system would not solve these questions by generating more scenarios. It would help identify where formal classification stops settling the theological problem.

From verification islands to a formalization frontier

This led me to think of theology as containing regions with different verification characteristics.

Some tasks permit relatively rapid checking: whether a quotation appears in a source, whether a canon states the claimed norm, whether terminology remains consistent, whether a conclusion follows from specified premises, whether a historical chronology is possible or whether two propositions contradict one another inside a defined corpus.

Other tasks remain partially verifiable: whether an interpretation is faithful to a whole tradition, whether a modern category distorts a historical text, which authority should govern a disputed question or whether an analogy illuminates more than it conceals.

Still other questions resist a simple external verifier: whether a person has encountered God, whether an interpretation is spiritually fruitful, what fidelity requires in one concrete life or how a community should discern an unprecedented situation.

I first described the highly structured regions as verification islands inside an interpretive ocean. The image was useful, but it risked making the boundary static. In practice, the boundary could move. Adding another source, historical period, tradition, language or pastoral context might turn one apparently closed question into several competing questions.

The more precise research object may therefore be a formalization frontier: the changing region at which theological operations become sufficiently bounded for rapid computation, and the point at which widening the model introduces kinds of meaning that its verifier cannot rank.

The system could identify the smallest premise whose modification changes an entire family of conclusions. It could compare Christianity, different Christian traditions, Buddhism or other religious systems without assuming that their inherited labels correspond to the deepest computational structures. It might discover families organized by authority, personhood, revelation, causality, liberation, ritual or soteriology that cut across conventional classifications.

Such results would be intriguing, but they would need careful interpretation. A computationally discovered cluster is not automatically a theological family. Similar formal structures can carry different historical and lived meanings. The model might reveal a relation that deserves investigation; it would not settle what the relation means.

Machine-scale theology created another boundary problem

The computational scale also changed my idea of what the research artifact might be. AI could generate argument structures too large for any person to read: millions of scenarios, networks of doctrinal implications, maps of disagreements, sensitivity analyses and complete histories of recursive revision.

This could make new forms of inquiry possible. A system might locate stable invariants across traditions, identify rare boundary cases, find recurrent contradictions or expose bifurcation points where one altered premise changes thousands of downstream conclusions.

My initial excitement focused on the possibility that such a study could not be performed manually at the same scale. Then another objection appeared. If no person can audit the whole structure, on what basis does it become theological knowledge rather than an enormous machine-produced object?

Machine-level verification could test consistency, provenance and reproducibility across the argument space. Human-level intelligibility would still require representative cases, traceable reasoning paths, summaries, boundary declarations and an account of why the result matters. A system may be computationally inspectable without being humanly comprehensible.

This prevents “the human role” from being defined as whatever operation remains inconvenient to automate this year. AI capability will continue moving. A more durable account locates human responsibility in choosing and revising boundaries, interpreting significance, authorizing sources, recognizing affected persons and deciding whether the formalized objective remains worth pursuing.

The method and the object began changing each other

Looking back, our conversation had used mathematical and engineering habits almost from the beginning: decomposition, constraints, failure modes, false-positive rates, sensitivity, multi-objective optimization, formal verification, stopping conditions and auditability.

Yet the process did not consist of placing technical vocabulary over theology. Engineering clarified the structure of the theological problem. Theology then exposed assumptions concealed by the engineering model.

Engineering asked:

Does the system satisfy its specification?

Theology answered:

Who wrote the specification? Which authority made its categories legitimate? Toward which good is the system ordered? Which persons, experiences and traditions became invisible so that verification could become quick?

This reciprocal correction is what makes the emerging method more than a superficial interdisciplinary combination. Without engineering, claims about AI and theology can remain impressionistic. Without theology and hermeneutics, formalization can mistake its chosen boundary for reality itself.

I would provisionally call the method boundary-relative computational theology. That name did not exist at the beginning of the conversation. It became possible only after several earlier answers failed in productive ways:

description of my method → evaluation of the method → doubt about the evaluation → exposure of hidden values → distinction between possibility and defect → problem of criticism without a null result → bounded completeness → computational theology

The final stage was not secretly contained in the first question. I could not have asked about boundary-relative computational theology before becoming dissatisfied with the apparently reasonable statement that my essays “may” suffer from certain defects.

What I now think and what I still do not know

I began by asking AI whether my way of thinking was good. The AI answered with strengths and weaknesses because that is what an apparently balanced evaluator is expected to do. My dissatisfaction with one part of the answer forced a distinction between possible failure and demonstrated defect. That distinction exposed the ambiguity of epistemic modality, the absence of a null result, the multiple objectives of humanistic evaluation and the open space of interpretive questions.

The investigation then turned back upon itself. If an AI can always produce another objection, what makes any one objection authoritative? If no neutral position ranks every legitimate theological value, what exactly is the evaluator optimizing? If completeness requires a bounded question, who establishes the boundary? And if AI makes millions of bounded theological operations possible, does it deepen theology or change the subject until only its verifiable residue remains?

I now think that AI criticism should be required to distinguish demonstrated error, material weakness, unmanaged risk, framework-dependent disagreement, possible extension and stylistic preference. It should identify its jurisdiction, show textual evidence, explain the consequence of leaving the passage unchanged and disclose what its proposed revision might sacrifice. Above all, it must be allowed to conclude that no material defect has been demonstrated.

I also think AI’s theological capability should be represented as a profile across bounded operations rather than one claim that it can or cannot “do theology.” The boundary should be recorded, varied and audited. Internal verification should never be confused with validation of the boundary itself.

Several questions remain unresolved. Can a system help evaluate the adequacy of its own boundary without beginning an infinite regress? Who has authority to decide that a theological corpus is sufficiently representative? How should machine-scale findings be made intelligible without reducing them to a few human-readable anecdotes? Can an AI produce correct theological distinctions without participating in the formation through which those distinctions become wisdom? And what happens when several traditions define successful theological reasoning differently?

The strongest question is still the one that appeared only near the end:

What must theology become in order to be rapidly verifiable—and what ceases to be theology when that transformation goes too far?

Before AI, scarcity of time provided criticism with an accidental stopping condition. AI removes that practical boundary without supplying an epistemological boundary to replace it. The next task is therefore not simply to make AI more critical. It is to construct conditions under which criticism can distinguish discovery from possibility, rank its own relevance, disclose its jurisdiction and legitimately stop.

Criticism is not self-validating. Completeness begins with a boundary. The boundary does not only limit what the system can know. Once made visible, it becomes one of the most revealing things the system allows us to study.

References

  1. Benzmüller, Christoph, and Bruno Woltzenlogel Paleo. 2014. “Formalization, Mechanization and Automation of Gödel’s Proof of God’s Existence.” Frontiers in Artificial Intelligence and Applications, vol. 263.
  2. Benjamini, Yoav, and Yosef Hochberg. 1995. “Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing.” Journal of the Royal Statistical Society Series B 57 (1): 289–300.
  3. Catholic Church. 1983. “Code of Canon Law, Book IV, Canons 959–997.” Vatican.
  4. Deb, Kalyanmoy, Amrit Pratap, Sameer Agarwal, and T. Meyarivan. 2002. “A Fast and Elitist Multiobjective Genetic Algorithm: NSGA-II.” IEEE Transactions on Evolutionary Computation 6 (2): 182–197.
  5. Eco, Umberto. 1990. The Limits of Interpretation. Bloomington: Indiana University Press.
  6. Hoare, C. A. R. 1969. “An Axiomatic Basis for Computer Programming.” Communications of the ACM 12 (10): 576–580.
  7. Mason, Elinor. 2023. “Value Pluralism.” Stanford Encyclopedia of Philosophy, substantive revision June 4, 2023.
  8. McAleese, Nat, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. 2024. “LLM Critics Help Catch LLM Bugs.” arXiv:2407.00215.
  9. OpenAI. n.d. “Working with Evals.” OpenAI API Documentation. Accessed August 21, 2026.
  10. Popper, Karl R. 1959. The Logic of Scientific Discovery. London: Hutchinson.
  11. Zheng, Lianmin, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.” Advances in Neural Information Processing Systems, Datasets and Benchmarks Track.