Archive777 isn't limited to conventional research questions. Some of the most interesting investigations begin with claims that are dismissed, disputed, forgotten, or labeled conspiracy theories.
Instead of deciding beforehand whether a claim is true or false, we ask a different question:
What do the source documents actually say?
That can mean searching declassified intelligence records, government reports, court documents, historical archives, scientific literature, congressional material, corporate records, or enormous public document releases.
The same approach can be used to investigate government misconduct, institutional failures, conflicts of interest, historical intelligence programs, corporate conduct, and allegations of corruption—following the evidence wherever the public record leads.
Sometimes the documents undermine a claim. Sometimes they add context that changes it. And sometimes official records confirm that something widely regarded as extraordinary really did occur.
Archive777 isn't built to prove a predetermined conclusion. A collection can support a claim, contradict it, complicate it, or leave the question unresolved. The point is to make the underlying evidence easier to find and examine.
The objective isn't to make the AI the authority. It's to find the evidence and let people examine it themselves.
Finding the documents is part of the research
A major part of Archive777 happens before anyone asks the AI a question.
Public does not necessarily mean easy to find.
Documents can be scattered across government websites, archives, FOIA reading rooms, court repositories, research databases, historical collections, public APIs, bulk releases, and poorly indexed legacy systems. Links disappear. File names can be meaningless. Collections may contain thousands of documents with very little context.
So we don't simply wait for documents to arrive and put an AI interface on top of them.
We actively look for hard-to-find public source material, acquire it, organize it into research collections, preserve useful provenance, and make it searchable.
We find the public files others don't.
More than 1.4 million files
Archive777 has grown to more than 1.4 million files across a wide range of research collections.
Those collections include subjects such as MK-ULTRA, the Epstein Files, Aliens & UFOs/UAP, 9/11, COVID-19, the Osama bin Laden Compound documents, historical intelligence programs, religious and historical texts, and many others.
Some collections contain thousands of files. Others contain tens of thousands, and some of the collections we're building are considerably larger.
At that scale, traditional research becomes difficult.
A researcher may know that an important document exists somewhere in a collection without knowing its filename, date, agency, folder, or the exact terminology used inside it.
That's where retrieval becomes useful.
See Archive777 in action
Here's a short demonstration of Archive777 researching its source-document collections:
Watch the Archive777 demo on YouTube:
https://www.youtube.com/watch?v=sNTHVSlu4q8
Collection building never really ends
A research collection isn't finished simply because we've downloaded a large number of files.
We continue looking for missing releases, related repositories, newly published records, historical versions, metadata, and source URLs. When agencies or archives publish new material, collections can grow again.
In other words, Archive777 isn't just a static document dump. The collections themselves are ongoing research projects.
From archive to RAG
The basic research pipeline looks something like this:
Find → Acquire → Organize → Preserve provenance → Index → Retrieve → Answer → Cite → Verify
Archive777 uses retrieval-augmented generation (RAG) so the model can research indexed source material rather than relying only on information learned during model training.
Our infrastructure includes self-hosted components such as AnythingLLM and Qdrant as part of the retrieval stack.
But the model is only one layer.
The quality of a research system depends heavily on the material underneath it: what was collected, where it came from, how it was organized, whether provenance was retained, and whether the researcher can get back to the source.
Provenance matters
Finding an interesting passage isn't enough.
Where possible, we preserve information that helps connect a file back to where it came from, including source URLs and collection metadata.
That becomes especially important when working with controversial subjects.
A generated answer saying that a government document contains something unusual isn't nearly as useful as being able to inspect the document yourself.
That is why Archive777 is being built around a simple principle:
Don't take the AI's word for it.
Follow the citations. Read the underlying documents. Check the source. Decide what the evidence means for yourself.
AI should help investigate evidence, not replace it
Large language models are extraordinarily useful research tools, but they can also be confidently wrong.
For Archive777, the answer isn't to pretend that problem doesn't exist.
The answer is to design the research workflow so the AI isn't the final authority.
The AI helps locate relevant material, connect information across large collections, and explain what it finds.
The citations and underlying source documents are where verification happens.
This distinction becomes particularly important when investigating disputed historical events, intelligence programs, allegations of misconduct or corruption, scientific controversies, and claims commonly described as conspiracy theories.
The question shouldn't be:
"Does the AI believe this?"
The better question is:
"What evidence exists in the source documents?"
Building this at scale is an ongoing engineering problem
Managing more than a million files creates problems that aren't obvious when experimenting with a few PDFs.
Duplicate documents matter.
Metadata matters.
Storage matters.
OCR quality matters.
Chunking and retrieval quality matter.
Vector database performance matters.
Source provenance matters.
Documents disappear from the web.
Government repositories change.
Collections receive new releases.
And sometimes the most valuable part of the work is simply discovering that a dataset or document collection exists in the first place.
Archive777 is still evolving as we work through those problems.
Why we are building it
There is already an extraordinary amount of information available in the public record.
Much of it simply isn't convenient to find or research.
We don't think the solution is asking people to trust another algorithm to tell them what is true.
We think we can build better tools for finding the records, searching them, connecting them, and putting the evidence in front of people.
That's what Archive777 is trying to do.
The AI searches. The sources provide the evidence. You decide what it means.
If you're interested in source-grounded AI research, large document collections, RAG, public records, archives, or investigative research, you can explore the project at:
Confirming Through Admittance.
Top comments (0)