Reviews run for a minimum of one week. The outcome of the review is decided on this date. This is the last day to make comments or ask questions about this review.
Eclipse INCEpTION
Eclipse INCEpTION is the successor to WebAnno, whose codebase it vendored and has continued to develop. WebAnno was developed at the UKP Lab and the Language Technology Group at Technische Universität Darmstadt with funding from CLARIN-DE; INCEpTION itself was developed at UKP Lab under a DFG grant.
Research funding is readily available for building new tools but rarely for maintaining existing ones, so as grant funding tapered the project's principal maintainer, Richard Eckart de Castilho, secured continued development through a series of freelance consulting contracts. That arrangement has kept INCEpTION actively developed and released, and it is expected to continue. UKP Lab, however, is now planning to wind down its engagement after many years of support. The project therefore needs a durable institutional home for the parts that still rest on UKP Lab resources, including the Maven repository and the demo server, and for the codebase itself.
A vendor-neutral foundation addresses this in two ways. It provides infrastructure and governance that do not depend on any single research institution's continued interest, and — in light of the EU Cyber Resilience Act — it can act as the open source steward for the codebase, taking on the obligations the Act places on stewards. Without a steward in that role, it is doubtful that the current maintainer and the project's occasional contributors could continue to work on INCEpTION as they do today.
Eclipse INCEpTION provides a platform for interactive and automated annotation of text and multi-modal content, supporting human-in-the-loop workflows, external machine-learning recommenders, and AI-assisted annotation capabilities.
Eclipse INCEpTION is currently mainly focussed on text, but is also already allowing certain multi-modal elements (e.g. linking text to images or displaying multi-modal contents embedded in HTML files or annotating PDF documents), so further extensions into the multi-modal space are possible.
Eclipse INCEpTION is mainly focussed on interactive annotation and human-in-the-loop scenarios, however, it also has started an experimental bulk processing module which can run annotation recommenders, auto-accept their suggestions and apply them to documents. In such scenarios, annotators might then use the annotation and quality assurance user interfaces mainly for inspection rather than annotation.
Machine-learning capabilities used in Eclipse INCEpTION are acquired via existing libraries (e.g. OpenNLP) or via external user-provided recommenders talking to Eclipse INCEpTION via an API. Development of machine learning capabilities within Eclipse INCEpTION is not planned. However, the ability to use LLM-based AI assistants e.g. in the form of a copilot chat is available as an experimental feature and planned to be developed further.
Eclipse INCEpTION offers a browser-based platform for semantic text annotation. Text annotation is the task of adding structured information to documents: marking which spans mention a person, a disease or a gene, recording how those mentions relate to each other, and linking each one to a concept in a knowledge base. The annotated corpora that result are the raw material for linguistic and digital humanities research, for domain studies in fields such as medicine and law, and for training and evaluating machine learning models. Producing them is slow manual work, usually spread across several annotators whose output has to be reconciled into a single reliable gold standard.
Eclipse INCEpTION covers that entire workflow in one application. Users define their own annotation scheme in the browser, ground it in their own ontology, annotate alone or as a team, and let machine learning recommenders that train on the work done so far take over the repetitive part.
- Annotation against an ontology. Load RDF, OWL, OBO, SKOS or Turtle, or query a remote SPARQL endpoint live. Ready-made profiles for Wikidata, SNOMED CT, the Gene Ontology, the Human Phenotype Ontology and GND.
- Arbitrarily layered schemes. Entities, relations, coreference, syntax, frames and document labels over the same text, with typed features and slots, all configurable in the browser.
- Suggestions that improve while you work. Recommenders train on the annotations made so far; active learning surfaces the cases they are least certain about. Nothing is written to the corpus until the annotator accepts it.
- Quality control. Curate multiple annotators into a gold standard, measure inter-annotator agreement, and inspect the collected data in the Explorer.
- Flexible deployment. A desktop installer for a single researcher, or a server deployment for an entire institution, on self-hosted hardware and integrated with the organisation's single sign-on.
- Programmatic access. A REST API, webhooks, and support for user-supplied models through the external recommender interface.
- Interoperable formats. Imports plain text, PDF, HTML and TEI; exports UIMA CAS XMI/JSON with custom layers intact, or CoNLL-U.
Eclipse INCEpTION has since become a staple of the text annotation community — it was the most-used annotation tool reported at LREC 2026, and more than 120 papers published in 2025 and 2026 alone have used it (https://inception-project.github.io/papers-using-inception/).
The term "INCEPTION" seems to have been registered by various parties as a trademark, including a the movie and an apparent recent filing under examination in the area of "computer products" (https://www.trademarkelite.com/europe/trademark/trademark-detail/019340196/INCEPTION) which would likely contend with a potential "Eclipse INCEpTION". Thus, it is expected that incubating the project within Eclipse would entail a rename including renaming packages within the codebase.
Eclipse INCEpTION currently uses the last version of Liquibase that was still under an OSI-approved license. It is not yet clear if, how and when it may be possible to switch this out against another approach. The high-profile Keycloak project has the same problem as is being watched for their strategy going forward.
The Eclipse Foundation is a good fit for several reasons. Its processes are well defined and consistently applied, and its release process scales down to a project with a small committer base: once a release has been approved, it can be executed by a single committer, without the quorum of votes that some other foundations require for every release. Its European base and its established engagement with the Cyber Resilience Act are also directly relevant to a project with European research roots and a largely European user and contributor community.
Eclipse INCEpTION has a backlog of issues to be worked on: https://github.com/inception-project/inception/issues
Strategically, current development directions are:
- better support for structured HTML/XML-based documents
- expanding functionality of the AI chat assistant
- expanding support for more ontologies and the ability to better work with them
An initial contribution could in principle happen any time once the formalities have been cleared.
If the project is accepted, it a refactoring of the code base will likely be required to account for the new name and home, e.g. renaming packages to start with org.eclipse.XXX wherever this is technically possible. This might entail some technical hurdles which could induce some delays or compromises (e.g. retaining old names in certain places for compatibility reasons). I would expect this refactoring to be completable in 4-8 weeks.
Typically, the release schedule of INCEpTION includes bug fix releases every two weeks and feature releases every 12 weeks (using a 2 digit versioning scheme). It is planned to continue this cadence once the migration is complete.
Prof. Dr. Iryna Gurevych (Technische Universität Darmstadt) (Co-PI on the already complered DFG project that gave rise to the code.)
Eclipse INCEpTION is presently released under the Apache License 2.0 with various bundled resources and dependencies under compatible licenses.
Large parts of the code were developed by people under contract at Technische Universität Darmstadt, so an explicit software grant from that party would be adequate. However, a lot of work was also contributed in spare time (as volunteer work) or within the scope of freelance consulting (with IP carveouts for the OSS contributions) as well as by external contributors under under Section 5 of the Apache License. From some contributors, we have had CLAs.
The project has vendored parts of various third-party codebases and developed them further, e.g. brat (https://brat.nlplab.org, MIT, unmaintained), pdfanno (https://github.com/paperai/pdfanno, github repository no longer exists, MIT), mtas (https://github.com/textexploration/mtas, AL, unmaintained). There is extensive documentation in the project about vendored and bundled third-party code and resources. Care was taken that third-party dependencies are compliant with the ASF 3rd Party License Policy (https://www.apache.org/legal/resolved.html) although INCEpTION additionally allows accepts LGPL dependencies.
A full list of the dependency and vendored code licenses is here. Note that many dependencies are multi-licensed, such that the listed license does not necessarily correspond to the license chosen for inclusion in INCEpTION (in particular dependencies that also list GPL in their licenses). A detailed list including assignment of the licenses to their respective libraries is available via the About page accessibly via the footer in an INCEpTION instance - e.g. on the demo instance at https://morbo.ukp.informatik.tu-darmstadt.de/about - this page is automatically generated by a ledger for vendored resources maintained in the codebase as well as by scanning declared license information in dependencies.
- Log in to post comments
Eclipse INCEpTION offers a browser-based platform for semantic text annotation. Text annotation is the task of adding structured information to documents: marking which spans mention a person, a disease or a gene, recording how those mentions relate to each other, and linking each one to a concept in a knowledge base. The annotated corpora that result are the raw material for linguistic and digital humanities research, for domain studies in fields such as medicine and law, and for training and evaluating machine learning models. Producing them is slow manual work, usually spread across several annotators whose output has to be reconciled into a single reliable gold standard.
Eclipse INCEpTION covers that entire workflow in one application. Users define their own annotation scheme in the browser, ground it in their own ontology, annotate alone or as a team, and let machine learning recommenders that train on the work done so far take over the repetitive part.
- Annotation against an ontology. Load RDF, OWL, OBO, SKOS or Turtle, or query a remote SPARQL endpoint live. Ready-made profiles for Wikidata, SNOMED CT, the Gene Ontology, the Human Phenotype Ontology and GND.
- Arbitrarily layered schemes. Entities, relations, coreference, syntax, frames and document labels over the same text, with typed features and slots, all configurable in the browser.
- Suggestions that improve while you work. Recommenders train on the annotations made so far; active learning surfaces the cases they are least certain about. Nothing is written to the corpus until the annotator accepts it.
- Quality control. Curate multiple annotators into a gold standard, measure inter-annotator agreement, and inspect the collected data in the Explorer.
- Flexible deployment. A desktop installer for a single researcher, or a server deployment for an entire institution, on self-hosted hardware and integrated with the organisation's single sign-on.
- Programmatic access. A REST API, webhooks, and support for user-supplied models through the external recommender interface.
- Interoperable formats. Imports plain text, PDF, HTML and TEI; exports UIMA CAS XMI/JSON with custom layers intact, or CoNLL-U.
Eclipse INCEpTION has since become a staple of the text annotation community — it was the most-used annotation tool reported at LREC 2026, and more than 120 papers published in 2025 and 2026 alone have used it (https://inception-project.github.io/papers-using-inception/).
- Log in to post comments
Support for bringing INCEpTION to Eclipse
Submitted by Florian Borchert on Fri, 09/25/2026 - 07:11
I strongly support the proposal to bring INCEpTION to the Eclipse Foundation. We used INCEpTION extensively in the GGPONC project, a German medical text corpus, which included the integration of custom ML models for annotation recommendations. The platform was very flexible and I was surprised how straightforward it was to adapt it to our project-specific needs.
It was also very easy to deploy and host on our university infrastructure, and the administration documentation was excellent. Whenever we ran into issues, bugs were addressed very quickly, and the frequent release cadence gave the project a strong sense of being actively maintained.
For projects that rely on annotation infrastructure over several years, that combination of flexibility, maintainability, and responsive development is extremely valuable. I would be very happy to see INCEpTION find a long-term home at Eclipse and think this would be a fantastic step for the project and its community.
An invaluable tool
Submitted by Andreas Zankl on Sun, 09/27/2026 - 05:58
I am a medical researcher at the University of Sydney. I run an Inception server to support my research. We use Inception to annotate clinical descriptions of rare diseases in the medical literature. The goal is to create a Gold Corpus of such annotations that can be used to train and evaluate LLMs to automate this task. Automated extraction of clinical descriptions from the medical literature would have many useful clinical applications. I evaluated many annotation tools for this task and only Inception fulfilled all my requirements. It has a great user interface, supports the most complex annotation schemes, handles large and complex medical ontologies, supports multiple annotators, has built-in recommenders, allows export in BioC format, the list goes on. In addition, Inception is well documented and Richard provides excellent support. With the rise of LLMs, a tool to manually curate text is more important than ever. AI is nothing without well annotated data. I wholeheartedly support this application to the Eclipse Foundation.
Interested party
Submitted by Monica Berti on Mon, 09/28/2026 - 04:57
Hello, as an interested party, I strongly support this project. I have used INCEpTION extensively at my institution, and it has become an essential tool in my research and teaching.
Hello, I am researcher at a…
Submitted by Marketa Preininger on Wed, 09/30/2026 - 01:50
Hello, I am researcher at a university and our international team uses INCEpTION for tagging. It is a useful tool and I plan to submit a proposal for third party funding in which I will propose using INCEpTION because it is user-friendly and very useful at the same time.
Support to the proposal from CLARIN Italy
Submitted by Francesca Frontini on Thu, 10/01/2026 - 03:52
We strongly support the proposal to establish INCEpTION as a new project at the Eclipse Foundation to ensure its long-term maintenance and development.
As an institution offering INCEpTION as a service to our research community within CLARIN-IT, we have relied on INCEpTION as a robust platform that integrates exceptionally well with our infrastructure (such as our Vocabularies and Thesauri publishing services).
Currently, we maintain our service across two instances—our main instance and a dedicated project instance—totalling 58 projects and at least 39 registered users from multiple academic and research institutions.
The INCEpTION developers are always very reactive in their support and have greatly helped us make our service accessible via CLARIN Single Sign-On, which makes our instance easily usable by academic users. We are currently in a beta testing phase, but after the migration to a larger server, we plan to offer INCEpTION to a broader community. In particular, there is great demand for classroom use, where lecturers are less likely to set up a local instance.
We would like to be listed as an interested party representing CLARIN-IT / CNR-ILC.
I am writing on behalf of…
Submitted by Maria Gavriilidou on Mon, 10/05/2026 - 06:40
I am writing on behalf of CLARIN:EL, which has offered an INCEpTION instance as a service to our research community for 7 years now. We fully support the proposal to move INCEpTION to the Eclipse Foundation, as we believe it will help secure the long-term maintenance and development of a tool that many of us rely on.
Our instance currently serves 134 users working in active annotation projects. They come from multiple countries across Europe and beyond, and they access the platform either through institutional accounts with federated login or through local accounts created in our repository. At present we host 117 active annotation projects containing 13,004 active annotations.
The platform has been used in a wide spectrum of activities from our part:
Given how broadly INCEpTION is used in our community, from education to advanced research, a sustainable, well-governed future for the project matters a great deal to us. We would be glad to be listed as an interested party in the proposal.
A very complete fully open-source annotation platform
Submitted by Grégoire Montcheuil on Wed, 10/07/2026 - 11:22
My team has been using INCEpTION for several years on in-house research projects to annotate medical documents. Over that time, we have observed the platform's rapid evolution.
We recently carried out a comparison of several annotation platforms. INCEpTION proved to be one of the most complete overall and, among fully open-source solutions, the richest in functionality.
Where most open-source annotation tools cover only part of the workflow, INCEpTION offers an exceptionally rich set of functionalities within a single application. It brings together, among other things:
And the user keep the ability to switch between several annotation interfaces.
We agree with some improvement areas identified in the proposal: a better support for structured and layout-preserving formats, and richer interfaces to work with ontologies.