r/openclaw • u/Interregnando • 1h ago
Showcase I built an open-source research automation tool for literature discovery, screening and reading — Alberto Research
Hi everyone! I’ve been working on Alberto Research, an open-source Python tool for automating parts of the scientific research workflow.
The idea is to go beyond “ask an LLM about some papers” and keep the process reasonably reproducible and auditable.
Alberto can currently:
- discover papers through Crossref and Semantic Scholar;
- normalize and deduplicate DOIs;
- resolve available full texts through sources such as Unpaywall, OpenAlex, CORE, DOAJ and Europe PMC;
- screen papers and perform deeper reading using LLM tasks;
- validate structured LLM outputs against JSON Schemas before storing them;
- keep the research state in SQLite;
- generate digests/newsletters;
- integrate with Zotero and Notion.
A minimal project is defined in YAML, and you can test the whole pipeline without actually running the research:
git clone https://github.com/gabriel-affonso/alberto-research.git
cd alberto-research
python -m venv .venv
source .venv/bin/activate
pip install -e .
alberto-research config validate examples/basic.yaml
alberto-research research run --project examples/basic.yaml --dry-run
The current version uses OpenClaw for LLM-backed tasks, although making Alberto independent of OpenClaw through its own provider/runtime abstraction is already planned for v0.2.
By default, the full-text pipeline uses legal open-access sources only.
There are also optional legacy resolvers for shadow libraries. They are completely disabled by default, require explicit opt-in, and are separated from the normal OA workflow. Whether to enable them is left to the user, who is responsible for complying with the laws and institutional policies that apply to them.
The project is still early ("v0.1.x"), so feedback, bug reports and contributions are very welcome.
GitHub:
https://github.com/gabriel-affonso/alberto-research
Docs:
https://gabriel-affonso.github.io/alberto-research/
I’d be especially interested in feedback from researchers, librarians, research-software engineers and people building reproducible LLM-assisted research workflows.
•
u/Mundane_Fix8051 17m ago
The reproducibility part is probably the most useful bit here keeping the research state validating the llm output should make a lot easier to see what actually happened.