How to know if your dissertation topic has already been done
From the Vermont Synergy Initiative’s methodology drafts (version 0.4, September 2026), word for word.
Draft for review · 31 August 2026 · not graded, not published Research room. Companion to How do you know that's true? (Methodology). Built to the VSI Research Translation Standard — three-tier keywords, 2 a.m. test.
Meta description (156 chars): Worried your research idea has already been done? One test tells an unasked question from a solved problem you couldn't find — and shows you where the open ground is.
Changelog v0.4. The real questions moved from the appendix to the opening: they are the reason this page exists, not evidence that the keywords are correct. New section, How this page was made, records the method applied to itself — including what it got wrong twice. Language brought to the Clarity Rule (below), which is a single standard for every reader and not an accommodation for any group.
The Clarity Rule, from Fred's Style Book:
"All content must be expressible at approximately a 9th-grade reading level without loss of meaning. Specialized terms are allowed only when defined or unavoidable. Failure to express an idea plainly signals incomplete understanding and requires revision."
This is not a simplified version written for some readers. It is the only version. Fred, 14 June 2026: "I speak to the intellectual elite in the field, and to newbies. From my side, the language is the same." And 18 July: "complexity is a way to hide ignorance... That's Feynman. If physics is too complicated to understand, that means you don't understand it either."
Jargon and idiom are barriers to everyone, including specialists — they slow reading and add work for the expert too. A reader with basic English is not a lesser audience. They are the test of whether the writing is actually clear.
The questions people actually ask
These are real. Each was posted in public, on a research forum, and each is quoted exactly as written.
"I am a phd student and need to start a research. I neither know where to start, nor how to choose a topic for my research. could you please recommend?" — ResearchGate
"i want to do a research paper for publication but i have an doubt there are many research papers are already exist from my research topic so how i know my solution or design is already existing or not" — Careers360
"Research already done by someone else — ever happened to you?" — FindAPhD Postgraduate Forum
"I don't know how to write a dissertation and i really need help!" — The Student Room
There are two different fears in those four sentences. One group has an idea and is afraid that someone else has already had it. The other group has no idea at all, and does not know where an idea is supposed to come from.
Both groups share one thing. Each has a great deal at stake and no instrument. They are asking strangers, in public, because nobody gave them a method.
This page is a method. It will not give you a topic. It will give you a way to find one, and a way to test whether what you found is real.
The short answer
Your fear is that your idea has already been done. That fear is reasonable, and it is the wrong thing to check first.
The larger risk is the opposite one: a question that looks open because nothing has ever tested it, and that is open for a reason you have not yet found.
So the first test is not has someone done this. The first test is:
What would have to be observed for this claim to be wrong?
If you cannot name anything — if no possible result would embarrass the claim — then you have found a question that nobody has actually asked. That is where a thesis can be built.
The rest of this page is how to tell that apart from the thing it closely resembles: a solved problem that you were not able to find.
Why you cannot tell them apart yet
You must choose a question and then live with it for three years.
Between you and that question sits a year of reading. The reading is not the hard part. The hard part is that a real gap and a closed one look identical from the outside. You will not learn which one you chose until you are deep enough that changing costs you a year.
Nobody gives you an instrument for this. You are told to read widely and to trust that a question will emerge. For some people it does.
The seven questions, and the inversion
A companion page in the Methodology room sets out seven questions for checking whether a claim is true. Point those same seven questions at a published literature instead of at your own notes, and they change purpose. Every failed answer becomes an opening.
| the question | what a failed answer means |
|---|---|
| Does it exist as cited? | a chain of papers citing each other where nobody read the original |
| Who wrote it? | a review citing a review — the first claim may not exist |
| Is it a whole thought? | a claim that appears only in abstracts |
| What question produced it? | a study cited for something it was never designed to test |
| What did the field conclude? | the finding, compared with what happened to it later |
| Has it been corrected? | a retraction that nobody downstream noticed |
| What does it forbid? | if nothing — that is unclaimed ground |
Ask them in order. You reach each question by passing the one before it.
Question one passes about twenty-nine times out of thirty. That is worth stating clearly. Does the quote exist where it says it does is the check that nearly everyone runs, and it is the weakest of the seven. It is necessary, and it proves very little. The value is in questions four to seven, and you reach those only by not stopping at the first.
The question that produces topics
A claim can collect a thousand citations without ever being tested in a way that could have failed. Citation counts measure usefulness. They do not measure exposure to risk.
An example from a real literature
EBV triggers MS.
This claim appears across many years of published work. The word "triggers" is rarely defined.
Does it mean that the virus starts the disease and then plays no further part? Or does it mean that the virus is still driving the disease, so that removing the virus should help?
Those describe two different diseases. They require different treatments and different trials. The word covers both meanings and commits to neither, and a claim that covers both commits to nothing.
Unverified — Grade U. A stronger form of this example — that a trial was run whose result contradicted the meaning used to justify it — comes from an earlier draft and has not been checked against primary sources. Do not cite it. It is kept here because it can be tested, and testing it is an exercise in the method below. If it turns out to be wrong, that belongs on this page too.
The pattern you can reuse: find a causal word doing precise mechanical work inside a claim while being used loosely. Triggers. Drives. Mediates. Modulates. Contributes to. Where the mechanism word is vague and the conclusion drawn from it is precise, the space between them has not been examined.
Why the openings are systematic, not luck
The published record is not a sample of what is true. It is a sample of what passed through five filters.
- Funding decides what is studied.
- Design decides what is measured.
- Publication bias removes the negative results — a study that finds no effect is much less likely to be published than one that finds an effect, so the record over-represents things that appeared to work.
- Indexing decides what can be discovered.
- Ranking decides what appears first.
Each filter leans the same way — toward the mainstream, the commercial, the Western, the positive. They combine, and they combine in one direction. So the unexamined ground is not scattered at random. It has an address.
This description is ours and carries Grade U. It has not been checked against primary sources and should not be cited as established. It is stated here so that you know it is a hypothesis — and because testing it properly is itself a thesis.
If you know the filters, you know where to look: the regional journal, the negative result, the study nobody repeated to check, and the paper whose relevance is structural rather than topical — the one whose subject is wrong and whose method is exactly right.
The part that will save you a year
Most unexamined ground is unexamined for a good reason.
Attention is limited. Researchers have strong reasons to build on anything genuinely useful. The difference in how often papers are cited reflects real differences in quality more often than not. A field that considered something and moved on is usually correct. So if you assume that every gap you find is an oversight, most of the time you will be wrong — and you will not find out until late.
A gap is a hypothesis about the field, and it is usually wrong.
So the question is never is this unexamined. A great deal is unexamined. The question is why, and there are only six answers:
| why it is unexamined | what it means for you |
|---|---|
| the field considered it and moved on | usually correct — choose something else |
| it could not be tested when last considered | check whether that is still true |
| it falls between two fields | nobody owns it, so nobody funded it |
| the method it needed did not exist | now it may exist |
| it is hard to fund, not uninteresting | possible, but plan the funding first |
| it was tested, it failed, nobody published | you will repeat it unless you find that result |
Only rows two to five are opportunities. Row one is the default, and you should assume it until you can rule it out. Row six is the trap: an unpublished failure is invisible by design, which is exactly why publication bias appears in the filter list above.
How to separate them, in one move: find the last person who looked. If a serious researcher examined it within the last ten years and moved on, that is row one, and you should believe them. If the last serious examination came before a method you now have, that is row four. Row four is where a thesis can be built.
The sentence that makes your claim defensible
A tool that finds gaps quickly will also find false gaps quickly.
"Nobody has asked this" is an absence being reported as a finding. An absence is usually a fact about your search, not a fact about the world. So the output of this process is never nothing is there. It is:
Here is what I searched, how, when, and with which terms — and within that, here is what I did not find.
That sentence will hold up in front of a committee. The other one will not. It is also your methods section, written a year early.
A failure, published deliberately
The pattern described above was built into an automatic detector and run over a 300 MB collection of documents, to see whether it could find another word like "trigger" on its own.
The result was negative.
| word | uses | defined | rate |
|---|---|---|---|
| produce | 20,016 | 427 | 2.1% |
| cause | 15,061 | 271 | 1.8% |
| trigger | 9,969 | 237 | 2.4% |
| affect | 9,262 | 165 | 1.8% |
| mediate | 2,285 | 34 | 1.5% |
Every causal word falls between 1.5% and 3.5%. "Trigger" is ordinary — slightly better defined than average. The detector could not separate it from "produce," so it would never have found the thing that made the pattern worth noticing.
The reason it failed is clear afterwards, which is the repeated lesson of this method. Causal words are almost never defined in ordinary writing, and they should not be. The sentence the rain caused the flood needs no definition of "cause." What made "trigger" a problem was register — an everyday word carrying precise mechanical weight inside a scientific claim. Register is a property of the sentence. It is not a frequency you can count.
The finding was made by reading. The attempt to automate it failed. Both appear here, because a method that reports only its successes is not a method.
What you actually do, this week
- Choose a claim your field repeats. Not a disputed one — a settled one.
- Ask what it forbids. If you cannot name an observation that would embarrass it, stop there and stay there.
- Find the last person who looked. Read them properly. If they considered it and moved on, believe them, and return to step one.
- Ask what has changed since then. A method, an instrument, a dataset, a population, a cost. If nothing has changed, it is probably row one.
- Write down where you searched and how — before you become attached to the answer.
- Check for retractions. This step protects your career rather than your argument. A retraction is almost never as loud as the claim it corrects.
Steps one to four are triage. Triage means sorting quickly to decide what deserves close attention, before you spend that attention. Triage compresses. What takes months of reading can be narrowed to hours by matching patterns across a large collection of documents.
Judgement does not compress. Whether a gap matters, whether it can be done, and whether it is worth three years of your life — that remains yours, and no amount of computing changes it.
We will not hand you a thesis. We will hand you the bench.
How this page was made
This section exists because the page teaches a method, and the honest way to demonstrate a method is to apply it to the page itself and report what happened — including the failures.
The search queries above were invented twice before they were found once. Versions 0.1 and 0.2 of this page listed the phrases we imagined a student would type. Our own Research Translation Standard names that exact error and calls it "the single most common failure mode." The Standard says plainly: "The operator does not invent these queries; the operator finds them."
Finding them changed the page. Reading what people had actually written produced two results that no amount of reasoning had produced.
First, the anxiety runs in the opposite direction from the one we assumed. We wrote a page about finding something. People arrive frightened of duplicating something. The opening was rewritten to answer the question that is actually being asked.
Second, an entire population was invisible to us. On ResearchGate, the topic-selection questions are overwhelmingly written by researchers working in English as a second language — eight out of eight results in one search. That is not a sampling accident. It is the shape of that community, and it is a group this page was ignoring while arguing that the literature is filtered toward the Western and the mainstream.
What is still missing, stated rather than hidden. The Standard names five sources for real queries. One was reached: public forum and page titles. Google's autocomplete suggestions and the "People also ask" panels were not reachable from the machine this was written on. So the list in the appendix is genuine and incomplete, and it is labelled that way.
That is the method performed on itself: check the claim, find the source, report what you could not reach, and correct what you got wrong rather than quietly replacing it.
What this is not
Not new in its parts. Citation checking, retraction checking and falsifiability are long established. What is offered here is the order, the inversion from verification to discovery, and the discipline of recording your search.
Not a replacement for reading your field. It makes the reading purposeful sooner.
Not reliable. It will produce false positives. The negative result above is ours.
Status and invitation
Ungraded. The examples are ours to have gathered. The grades are not ours to assign. Nothing here has been checked against primary sources except where marked, and two passages carry Grade U.
If you use this and it finds you something, tell us what it found. If you use it and it wastes your time, tell us that instead. That is the more useful report, and it is the one nobody ever sends.
Appendix — three-tier keyword architecture
Per the VSI Research Translation Standard, Part 5. The observed questions that open this page are the primary Tier 3 evidence. They are placed at the top of the document because they are the reason it exists, not because they are keyword data.
Tier 1 — Technical (specialist registry) research gap identification · falsifiability criterion · citation chain analysis · publication bias · null result reporting · retraction surveillance · literature saturation · thesis problematisation
Tier 2 — Accessible (bridge registry) how to find a research gap · what is a research gap · choosing a research question · how to tell if research is original · identifying gaps in the literature
Warning — this tier is saturated. "How to find a research gap" and "what is a research gap" are already served by at least eight commercial pages (Grad Coach, ResearchRabbit, Scispace, Enago, Researcher.Life, Editage, IU LibGuides and others). High volume, heavy competition, and the existing pages answer the question well enough. Do not build this page around that term.
Tier 3a — observed, Anglophone forums (captured 31 Aug 2026)
| observed phrasing | where |
|---|---|
| "Research already done by someone else — ever happened to you?" | FindAPhD Postgraduate Forum |
| "Discovering research has been done before?" | The GradCafe Forums |
| "Is there a way to check if a dissertation or thesis has already been written on a specific topic before starting to write it?" | Quora |
| "How to check if a research project is already ongoing somewhere in the world?" | ResearchGate |
| "How can you find out if your research paper topic has been done before" | Quora |
| "I don't know how to write a dissertation and i really need help!" | The Student Room |
| "Don't know where to start?" | Editage |
Tier 3b — observed, English as a second language
Fred Schwacke, 31 August 2026, correcting an earlier draft that had called this register "rough":
"Not so much rough as it appears to be ESL. And that will be a large and fertile group for us to address."
"I have an doubt" and "already existing or not" are not errors. They are Indian English constructions, and Careers360 is an Indian education portal. A researcher searching in a second language is demonstrating more competence than a native speaker searching in their first, not less.
This is a search-language finding, not a writing instruction. These phrasings tell us what to match, so the page can be found. They do not mean the page should be written differently for these readers. The page is written to the Clarity Rule, which is one standard for everyone. Fred, 31 August, making the distinction:
"It's not about writing to the ESL reader. It is about writing clear, simple, understandable language that anybody with rudimentary skills in the language can understand fully."
| observed phrasing | where |
|---|---|
| "Kindly suggest some topics for PhD Thesis (from English Literature and Cultural Studies)" | ResearchGate |
| "Would you kindly suggest me some research topics?" | ResearchGate |
| "I am a phd student and need to start a research. I neither know where to start, nor how to choose a topic for my research. could you please recommend?" | ResearchGate |
| "Kindly suggest me current research topic in the field of Remote sensing and GIS for Ph.D?" | ResearchGate |
| "Reasearch topic suggestion" | ResearchGate |
| "PhD research Topics [NEED GUIDANCE]?" | ResearchGate |
| "i want to do a research paper for publication but i have an doubt there are many research papers are already exist from my research topic so how i know my solution or design is already existing or not" | Careers360 |
Search markers for this register: kindly suggest · suggest me · need to start a research · I neither know · doubt (meaning question) · already existing or not · please recommend · NEED GUIDANCE.
The two populations want different things
| what they ask | what they have | how well served | |
|---|---|---|---|
| Anglophone forums | has this been done? | an idea, and a fear | adequately — many pages exist |
| ResearchGate / ESL | kindly suggest me a topic | no idea, and no instrument | barely at all |
The second group is larger and worse served, and it asks for something this project will not provide: a topic. The honest answer is we will not hand you a topic; we will hand you the instrument that produces one, and it works in any field. Stated plainly, that is a better answer than a list of suggestions, and it is the only one that respects the person asking.
2 a.m. test
| test | v0.1 | v0.2 | v0.3 | v0.4 |
|---|---|---|---|---|
| Surfaces to the query | ✗ | ~ | ✓ | ✓ |
| Answers without preamble | ✗ | ✓ | ✓ | ✓ |
| Gives the next concrete step | ✓ | ✓ | ✓ | ✓ |
| Tier 3 observed, not invented | ✗ | ✗ | ~ forums only | ~ forums only |
| Meets the Clarity Rule | ✗ | ✗ | ✗ | ✓ |
The last row is not "readable by an ESL researcher." That framing was corrected on 31 August: writing plainly is one standard for every reader, not a special version for one group. A reader with basic English is how you find out whether you met it.