Using an LLM to read and interrogate a doctoral thesis

Dr Haider Ali

This is a proposed workflow, not an evaluated one. It has not been tested against outcomes, and Proposition 6 in particular is conjecture. A note on how the article was produced appears at the end.

The most useful question is not "What do you think of my chapter?"

For students, the challenge of AI usually presents itself as a problem of writing. Large language models produce text that sounds seductively plausible, and using that text can create serious difficulties with an institution.

Much less widely discussed is the opposite capability: the ability of an LLM to read a student’s own material and surface issues that the student can then address themselves.

This matters most for long documents. A close reading of a multi-chapter thesis is time-consuming and genuinely difficult, and it becomes harder the longer the author has lived inside the manuscript. It is often easier for an external examiner encountering the text for the first time to notice that a promise made in Chapter 3 is never quite fulfilled. The author cannot easily see this, because the author knows how the argument is supposed to work and reads every ambiguous passage in the light of that knowledge.

An LLM does not possess that internal reconstruction either. Used carefully, that is precisely its value. But its value depends heavily on the task it is given. A broad request such as:

Read this chapter and tell me whether the argument is coherent.

may produce fluent and apparently insightful feedback, but it creates two problems.

  1. The judgement is difficult to audit. The student receives an interpretation without necessarily being able to see what evidence produced it.

  2. The model is easily drawn into the framing supplied by the user. Asking whether a chapter is coherent can encourage it to construct a coherent reading. Asking whether the contribution is convincing can encourage it to explain why the contribution is convincing.

The more productive use of an LLM is therefore not primarily as a substitute supervisor, examiner or expert reader.

It is as a systematic auditor of the thesis.

The distinction matters. A reader gives an impression. An auditor produces evidence that can be checked. Used in this way, the LLM can help expose patterns that are particularly difficult to detect after years of immersion in the same manuscript: definitional drift, unfulfilled promises, long-range contradictions, unsupported conclusions, changing assumptions and contributions that are asserted more strongly than they are demonstrated.

Seven propositions set out how to do this. A short section on institutional and ethical constraints comes first, because those constraints determine whether any of it is available to a given student at all.

Before uploading anything

This approach assumes that the student uploads substantial portions of an unexamined thesis to an LLM. That raises four separate questions, and answering one of them does not answer the others.

1. Where does the data go?

At the Open University, students may upload their work provided they are logged into Microsoft Copilot with their student account and the green shield indicating Enterprise Data Protection is visible. That means the material is not retained and is not used for training. Students at other institutions should consult their own university's policies rather than assuming an equivalent arrangement exists; consumer and enterprise deployments of the same underlying model can differ materially in how they handle uploaded content.

2. Does the ethics approval extend this far?

A thesis containing interview transcripts, survey responses or other participant data was collected under an approval, and under a participant information sheet that told those participants how their data would be processed. Enterprise Data Protection addresses the security of processing. It does not retrospectively widen the consent that was given. Before uploading empirical chapters, check what the approval and the information sheet actually say, and consider whether the audit can be run on anonymised or redacted versions of the data-bearing sections.

3. Is any of the material subject to a third-party constraint?

Industrially sponsored and collaborative studentships, non-disclosure agreements, commercially sensitive data and embargoed chapters all sit outside the institutional question entirely. A tenant boundary does not override a contractual one.

4. Does this use need to be declared?

Whether a use is permitted and whether it must be declared in the final submission are separate matters, and doctoral regulations on AI declaration are changing quickly and vary considerably between institutions. It is entirely possible to use an LLM in a way the university permits and still encounter difficulty for not having declared it. Establish the declaration requirement before starting the work, not after finishing it, and keep a record of what was done as you go — the audit trail is much easier to build contemporaneously than to reconstruct.


Proposition 1: Extract the argument before asking the LLM to evaluate it

The first pass through a chapter should usually not be a summary.

A summary can make an argument appear coherent precisely because it removes the small inconsistencies, qualifications and shifts in terminology that need to be examined. Summarisation is useful for orientation and potentially damaging for diagnosis.

A stronger starting point is to ask the LLM to extract structured artefacts from the chapter.

The essential difference between extraction and summary is not that one compresses and the other does not. Extraction compresses too, and a well-formed table can lend an appearance of clean architecture to prose that was in fact rather fluid.

The difference is that extraction can be indexical and a summary cannot. Every row should carry the relevant wording and a location — section, page, paragraph — so that the student can return to the manuscript and check it. Where wording matters, particularly for definitions and central claims, the LLM should reproduce the wording rather than paraphrase it.

The purpose is not to create a substitute version of the chapter. It is to make its argumentative architecture visible, and checkable. Once the material is in structured form, the student can ask questions that are difficult to answer reliably from continuous prose:

  • Is the same construct defined differently in different places?

  • Does a conclusion depend upon an assumption that was never defended?

  • Is an important qualification present in the literature review but absent from the discussion?

  • Does the chapter claim more than the evidence listed beneath the claim appears to establish?

  • Are several apparently different propositions actually repetitions of the same underlying claim?

The important principle is:

Extract first; evaluate second.

The LLM's initial job is to make the structure inspectable. Judgement should follow from the resulting evidence.

Proposition 2: Turn every important promise into a ledger — and treat the ledger as incomplete

Doctoral theses contain many forms of promissory statement.

Some are obvious:

  • Chapter 6 will demonstrate…

  • The empirical analysis will establish…

Others are embedded in conventional doctoral architecture: research questions, research objectives, hypotheses or propositions, methodological commitments, theoretical claims, promised analyses, stated limitations, claims of originality and contribution statements.

These should not merely be noticed. They can be audited systematically.

For each commitment, construct a ledger:

Human readers naturally remember the broad trajectory of an argument while forgetting individual promises made many pages earlier. A machine can search repeatedly for the explicit textual traces, and can do so consistently across a long document.

But this is the point in the process where the method is most likely to mislead, and it is worth being precise about why.

It is tempting to assume that an LLM is well suited to this task because the work is repetitive rather than intellectually demanding. That assumption is close to backwards. Current models are strong at fluent evaluative judgement and comparatively weak at exhaustive enumeration across long documents. Recall degrades as context lengthens, and it degrades quietly: the model does not report that it has stopped finding things.

The consequence is an asymmetry in failure modes that determines how the output should be read.

A false positive — a flagged commitment that was in fact discharged — is visible and self-correcting. The student checks the passages, sees the discharge, and dismisses the flag. The cost is a few minutes.

A false negative — a commitment that the ledger never lists at all — is invisible. There is nothing to check. And the natural reading of a ledger that comes back clean is reassurance, which is the one conclusion this method cannot support.

Two practical consequences follow.

First, enumerate the commitments by hand wherever the list is bounded. Research questions, objectives, hypotheses and stated contributions are almost always short, explicit and easy to list manually from the front matter and introduction. Compile that list yourself, then give it to the LLM and ask it to find the discharge. This uses the machine for the expensive, repetitive search across the manuscript, and the human for the enumeration where completeness matters most.

Second, spot-check the ledger's recall. Take three or four promissory statements you already know exist, buried in the middle chapters, and check whether the ledger caught them. If it missed one, that is information about how much weight the rest of the ledger can bear.

The ledger is a hypothesis generator. It is not a clearance certificate.

Where a commitment does appear undischarged, that is not automatically a thesis problem. The claim may have been reformulated, combined with another objective, or made redundant by a later analytical decision. But the discrepancy is worth seeing.

The purpose of the LLM is therefore to identify:

"Here is a commitment that I cannot find being clearly discharged."

It should not conclude:

"This commitment has not been discharged."

And the student should not read silence as:

"All commitments have been discharged."

That distinction — between locating candidate problems and deciding whether they are genuine, and between absence of evidence and evidence of absence — is fundamental to responsible use throughout the thesis.

Proposition 3: Read the thesis backwards from what it finally claims

Most students construct and repeatedly read their thesis forwards. Introduction leads to literature review; literature review to methodology; methodology to findings; findings to discussion; discussion to conclusion. This creates a powerful sense of narrative continuity. Each chapter seems to grow naturally from the preceding one. But coherence felt while moving forwards is not the same as argumentative sufficiency. A particularly valuable LLM pass therefore runs in the opposite direction. Begin with the conclusion. For every substantial concluding claim, ask:

What would have needed to be established earlier in the thesis for this conclusion to be warranted?

Then work backwards. For example:

Conclusion: The study demonstrates that X is an important mechanism explaining Y.

The backwards audit asks:

  • Where was X defined?

  • What evidence established that X occurred?

  • What evidence connected X with Y?

  • What alternative explanations were considered?

  • What analytical step justified describing X as a mechanism rather than merely an association?

  • Where were relevant limitations addressed?

  • Does the wording of the conclusion exceed the strength of the earlier evidence?

This process can be repeated for each major conclusion and contribution. The result is a claim-to-foundation map.

That map exposes a common doctoral weakness: a thesis may contain all of the relevant ingredients without demonstrating the logical connection among them strongly enough to support what is eventually claimed.

Backward reading therefore asks a more demanding question than:

Does the thesis tell a coherent story?

It asks:

Has the thesis earned the right to make its final claims?

That is much closer to the kind of question an examiner is likely to ask.

Proposition 4: Compare distant parts of the thesis directly

Many inconsistencies develop too gradually to be noticed through ordinary sequential reading.

Consider a construct introduced in Chapter 2. Its definition changes slightly in Chapter 3. The empirical operationalisation in Chapter 4 introduces another small shift. The findings chapter uses the terminology somewhat differently again. By Chapter 7, the construct may no longer mean exactly what it meant when the theoretical argument was developed. Every individual transition appears reasonable. The cumulative movement may not be.

This is why non-adjacent comparison is particularly valuable. Instead of asking only whether Chapter 2 is consistent with Chapter 3, and Chapter 3 with Chapter 4, ask the LLM to compare:

  • Chapter 1 directly with Chapter 7;

  • Chapter 2 directly with Chapter 6;

  • the literature review directly with the discussion;

  • the methodology directly with the claims made in the conclusion;

  • the original research questions directly with the final contribution statement.

These long-range comparisons are particularly useful for identifying:

  • Definitional drift. Has a key concept changed meaning?

  • Theoretical drift. Is the theoretical position used to interpret the findings the same one established in the literature review?

  • Methodological drift. Does the discussion make claims that the stated research design was actually capable of supporting?

  • Evidential drift. Have qualified findings become categorical claims by the conclusion?

  • Contribution drift. Is the contribution claimed at the end the same contribution the thesis set out to make?

This comparison should again be evidential. Rather than asking:

Has my definition of X changed?

ask:

Extract every substantive definition or characterisation of X from Chapters 2, 4 and 7. Place the passages side by side, with their locations. Identify differences in attributes, causal role, scope and boundary conditions. Do not decide whether the differences are problematic until after presenting the evidence.

One practical constraint deserves mention. A complete thesis may exceed what can be held in a single context, and if the chapters are processed separately then the direct comparison this proposition depends on has not actually taken place. Where the full text will not fit, extract the relevant passages first — using Proposition 1 — and run the comparison over the extracts rather than the chapters. That also makes the comparison auditable, since the extracts carry their locations.

The same recall caution from Proposition 2 applies here: "every substantive definition" is an instruction the model will attempt and may not complete. The LLM reveals the movement. The doctoral researcher interprets its significance.

Proposition 5: Ask the LLM to try to break the argument — and then discount its enthusiasm

One of the least useful ways to employ an LLM is to ask it repeatedly to confirm that the thesis makes sense. The more useful stance is adversarial.

Instead of Is this argument convincing? ask:

Where could a sceptical examiner challenge the inference being made here?

Instead of Is this chapter coherent? ask:

Identify the places where the argument is most vulnerable to the claim that the conclusion does not follow from the evidence presented.

Instead of Does the literature review justify my research gap? ask:

Assume an examiner believes the claimed research gap is overstated. Construct the strongest case for that position using only this chapter.

A doctoral student has usually spent several years making the strongest possible case for the thesis, and by the final stages knows how the argument is intended to work. An adversarial pass is useful precisely because the model is instructed not to repair the argument on the author's behalf. A particularly useful instruction is:

Do not infer a missing logical step merely because you can see how the author could have made it. Identify where that step needs to be explicit.

However — and this is the correction that matters — adversarial framing is not more truthful than confirmatory framing.

The reason a request for confirmation is unreliable is that the model tends to supply what the framing requests. That tendency does not disappear when the framing is reversed. Ask for the three weakest points in a chapter and you will receive three, whether or not three exist. Ask for the strongest case that the research gap is overstated and you will receive a case, whether or not the gap is overstated.

Adversarial prompting has not removed the bias. It has pointed it in a direction that is cheaper to check — an unwarranted criticism costs the student the time taken to reject it, whereas unwarranted reassurance costs a viva. That is a real advantage, and it is the only advantage. It is not a claim to accuracy.

Three practices follow.

  1. Require evidence with every criticism. Any objection that does not cite specific passages cannot be evaluated and should be set aside rather than argued with.

  2. Ask for calibration explicitly. Instruct the model to rank the objections by strength, to state which are substantive and which are presentational, and to say plainly where it considers an objection weak. Models will rarely volunteer that a chapter has no serious vulnerability, but they will often concede it when asked directly.

  3. Apply the four-question diagnostic set out at the end of this article before acting on anything. It exists for this pass in particular.

The same adversarial principle can be applied at several levels:

  • Sentence level. Does this inference follow from the previous statement?

  • Section level. Does the evidence presented justify the subsection conclusion?

  • Chapter level. Does the chapter accomplish the task assigned to it?

  • Thesis level. Does the complete argument support the contribution eventually claimed?

The LLM is not being asked to play the caricature of a hostile examiner. It is being used to create structured intellectual resistance — which is often more valuable than another agreeable reading, provided it is not mistaken for a verdict.

Proposition 6: Use repeated independent readings — but establish a baseline first

LLM outputs are not perfectly stable. For some tasks this is frustrating. For doctoral review, it can be diagnostically useful. Suppose the same question is posed three times in genuinely fresh contexts:

What is the central theoretical claim made in this chapter?

If all three readings independently identify essentially the same claim, that provides some evidence that the claim is clearly signalled. If the three readings identify substantially different claims, the useful question is not which answer is correct? but:

Why was this chapter capable of supporting three different readings?

Divergence becomes a possible signal of interpretive instability. If several machine readings diverge, human readers may also diverge.

The difficulty is that divergence has more than one cause, and the raw comparison cannot distinguish between them.

Output varies because of sampling behaviour, because of what else is in the context, because of the model version, and because of the exact wording of the prompt — not only because the manuscript is ambiguous. A perfectly clear chapter can produce three different "central theoretical claims" simply because central was never defined, or because the model sampled differently on each run. Treating all divergence as evidence about the thesis would be exactly the kind of unfalsifiable inference the rest of this article is concerned with catching. The proposition survives, but only with a control.

Hold everything constant except the manuscript. Same model, same version, same settings, same prompt wording, same absence of prior context. If any of those change between runs, the comparison is uninterpretable.

Specify the question tightly. What is the central theoretical claim? is underspecified and will generate divergence on its own. Identify the single claim this chapter asks the reader to accept, quoting the sentence in which it is most directly stated is answerable, and disagreement about it means something.

Establish a baseline. Run the identical procedure on a text you have independent reason to believe is clearly argued — a well-regarded published paper in your field, or a chapter your supervisor has already told you is unambiguous. Whatever divergence that produces is the method's noise floor. Only divergence above that floor is evidence about your own writing.

With the baseline in place, the student can classify results meaningfully:

  • Stable. Independent readings substantially agree.

  • Variable but compatible. Different readings emphasise different aspects without contradicting one another.

  • Materially divergent. Different readings imply significantly different understandings of the argument — and the divergence exceeds what the same procedure produced on the baseline text.

Only the third category warrants revision. Some theoretical arguments are appropriately nuanced or deliberately open; the point is not that ambiguity is always undesirable, but that the student should know where it exists and should not mistake instrumentation noise for it.

The method can be applied to the research gap, the meaning of a central construct, the relationship between two theories, the explanation of an empirical result, the thesis's principal contribution, or the logical purpose of a chapter.

Properly controlled, it produces an unusual but important principle:

Variation in AI output, measured against a baseline, can become evidence about the clarity of the manuscript.

Proposition 7: Audit the contribution separately from coherence

A thesis can be coherent without making a sufficient doctoral contribution. This deserves a separate pass.

Coherence asks whether the components of the thesis fit together. Contribution asks whether, having fitted them together, the thesis has produced something sufficiently original, significant and defensible.

Students should therefore extract every substantive contribution claim and build a contribution audit.

Each claimed contribution can then be interrogated.

  • Specificity. What exactly is new? Does the thesis identify a precise addition, correction, extension or challenge, or rely on vague formulations such as "adds to understanding"?

  • Contrast. New relative to what? Can the prior position in the literature be stated clearly enough to show what has changed?

  • Evidence. Which findings actually support the contribution? Could it survive if one major finding were removed?

  • Proportionality. Is the language of contribution proportionate to the evidence? Does the thesis demonstrate, suggest, extend, qualify, challenge or merely illustrate?

  • Location. Is the contribution visible only in the final chapter, or has the thesis prepared the reader to recognise why it matters?

  • Defence. What would the strongest sceptical response be? Could an examiner plausibly say that the result is interesting but already implicit in existing literature?

This audit is particularly valuable because contribution claims often become stronger during the final stages of writing. A finding initially described cautiously may gradually become a theoretical contribution through repeated revision. An LLM can compare the contribution language with the findings and theoretical argument from which it supposedly follows.

The objective is not to ask the model:

Is this contribution sufficiently original for a PhD?

That is ultimately a disciplinary judgement for appropriately qualified human examiners and supervisors. The more defensible use is:

Show me exactly what textual and evidential foundations this contribution appears to rest upon, and identify the strongest places where that chain could be challenged.

The division of labour

The seven propositions imply a particular relationship between doctoral researcher and LLM.

The LLM is useful for extraction, comparison, bookkeeping, searching, pattern detection, constructing counterarguments, tracing dependencies, identifying possible omissions, and repeating the same analytical procedure consistently across large amounts of text.

The student retains responsibility for deciding whether an apparent inconsistency matters, judging the quality of evidence, determining whether conceptual development is legitimate, deciding which criticism is disciplinarily important, interpreting ambiguity, assessing theoretical significance, deciding what the thesis ultimately claims, and revising the argument.

To that division, one qualification should be added, because it cuts across the whole of the first list.

Everything in the LLM's column is a task where the model produces candidates, not complete sets. It finds definitions but not necessarily every definition; it locates discharges but not necessarily every failure to discharge; it identifies drift where drift is textually obvious and misses it where it is subtle. What the model returns can be checked against the manuscript. What it fails to return leaves no trace.

The practical implication is that this method is well suited to finding problems and poorly suited to ruling them out. A student who runs all seven passes and finds little has not established that the thesis is sound. They have established that these particular procedures, on this occasion, did not surface anything — which is a much weaker statement, and should be recorded as such.

This is also why the most useful LLM outputs in doctoral work are not answers. They are candidate problems accompanied by evidence.

A good response therefore looks less like:

Chapter 4 is theoretically inconsistent with Chapter 2.

and more like:

Chapter 2 defines X using attributes A, B and C (Section 2.3, p. 41). In Chapter 4, X is operationalised using A and D (Section 4.2, p. 88), with no explicit explanation of the change from C to D. These passages appear to be the relevant evidence. Consider whether the operationalisation requires justification.

The second output is considerably more useful because the student can inspect it, disagree with it and act upon it.

A seven-pass audit

Pass 1: Argument extraction. For each chapter, extract claims, definitions, assumptions, evidence, qualifications, dependencies, forward references and contribution claims, with wording and locations. Do not summarise.

Pass 2: Commitment audit. Enumerate the bounded commitments — research questions, objectives, hypotheses, stated contributions — by hand. Use the LLM to locate their discharge. Spot-check recall before trusting the ledger.

Pass 3: Backwards audit. Begin with the final conclusions. For every major claim, identify what would need to have been established earlier, and trace each requirement back.

Pass 4: Long-range consistency audit. Compare non-adjacent chapters directly, working from extracts where the full text will not fit in a single context.

Pass 5: Examiner challenge. Ask for the strongest plausible challenge to the research gap, theoretical reasoning, methodological fit, evidential sufficiency, interpretation and contribution. Require textual evidence for each. Ask for the objections to be ranked and for weak ones to be identified as weak.

Pass 6: Independent-reading test. Establish a baseline on a text known to be clear, then repeat selected analyses on your own chapters under identical conditions. Treat only above-baseline divergence as diagnostic.

Pass 7: Contribution audit. Take each stated contribution separately and trace: claim → prior literature → evidence → reasoning → contribution.

Throughout, keep a record of what was run, on which chapters, with which model. If a declaration is required at submission, it is far easier to write from a contemporaneous log.

A final diagnostic

Before accepting any LLM-generated criticism, the student should ask four questions.

  1. What exactly has the model identified? Is this an inconsistency, an omission, an ambiguity, an unsupported inference, or simply an alternative interpretation?

  2. Where is the evidence? Can the relevant passages be located in the thesis? If not, the criticism cannot be evaluated.

  3. Could I explain why the issue matters? If not, the criticism may merely sound sophisticated.

  4. Would I make the same judgement after reading the passages myself? If not, the model has identified something to investigate, not something that must be changed.

And one question to ask of the whole exercise:

What would this method have missed?

The passes above search for particular kinds of failure — undischarged promises, definitional drift, unwarranted inference, overstated contribution. A thesis can be weak in ways none of them are looking for: a poorly chosen research question, a literature the author has not read, a method that was competently executed and inappropriate from the outset. None of those will appear in a ledger.

That final distinction is crucial. The optimal use of an LLM in doctoral revision is not to surrender judgement to a machine capable of reading the thesis quickly.

It is to use the machine's speed and capacity for repeated comparison to make the thesis more inspectable to its author, while remembering that what the machine does not report is not thereby absent.

The governing principle might therefore be expressed simply:

Do not ask the LLM to tell you whether your thesis is good. Ask it to expose the structure of your argument, trace what you have promised and delivered, search for where the reasoning may fail, and show you the evidence. Then make the judgement yourself — including the judgement about what it failed to find.

A note on how this article was produced, in the spirit of the run log it recommends. The approach began as a conversation with Claude about the underlying problem, was extended with material from Copilot, and was drafted into its proposition structure by ChatGPT. It was then revised again after applying something close to Proposition 5 to the article itself — which identified three substantive errors: an unexamined assumption that LLMs are reliable at exhaustive enumeration, a mirror-image bias in the adversarial pass, and a confound in the independent-reading test. A further gap surfaced only when the workflow was built as a tool, where the interface forced a question the prose had not: what to do when the baseline text itself produces divergent readings.

And, in the spirit of the limitation it recommends: none of this has been tested against outcomes. The propositions rest on documented model behaviour and on ordinary examining practice, not on evidence that following them produces better theses. Proposition 6 is the most speculative and should be treated as such. Readers may reasonably regard the article's own construction as either a demonstration of the method or a limitation of the piece.