Newsletter

Welcome!

Welcome to the July 2026 issue of the Text+ newsletter! This is a special edition dedicated to the 5th Text+ Plenary. In this issue, we would like to look back on the event, highlight key themes and ideas, and share the most important impressions and outcomes with you.

As usual, the newsletter is available on our website under News/Announcements; it is also sent by email to all interested parties. If you would also like to receive the newsletter by email, please do get in touch.

Between newsletter issues, we’ll also keep you up to date with the latest news via our social media channels. Please feel free to follow us on LinkedIn and/or Mastodon.

We’re happy to answer any questions, suggestions or feedback you may have. You can contact the Text+ Office at office@text-plus.org.

5th Text+ Plenary

On 15 and 16 June 2026, the 5th Text+ Plenary took place at Mannheim Palace. Invited guests included Text+ staff, representatives from other NFDI consortia and other interested parties from the community, including delegates from professional associations and research networks. This year’s Plenary was dedicated entirely to marking the conclusion of Text+’s first funding phase.

After nearly five years of collaborative work, the event provided an opportunity to look back on the results and developments achieved during the first funding phase, as well as to take a cautious look at the possible future of Text+. Through presentations, panel discussions and other formats for exchange, the milestones achieved were presented, and experiences from the task areas of Collections, Editions, Lexical Resources and Infrastructure & Operations were discussed, along with their potential prospects for the further development of Text+. The plenary session once again highlighted just how much the close collaboration within Text+ and the wider community has contributed to the consortium’s success, and how important ongoing exchange is for the further development of the research data infrastructure.

All presentations and posters presented during the public portion of the event are available on the Text+ Plenary 2026 event website.

Day 1 of the programme: 15 June 2026

Management Strategy Meeting

Even before the official start of the Plenary, on the morning of the first day of the event, the management team for a potential second phase of the Text+ project met for an internal strategy meeting. Subject to the approval of a second funding phase for the project, discussions centred on the future direction and further development of Text+, as well as plans for the transition from the first to the second funding phase.

Arrival and opening of the 5th Text+ Plenary

The public part of the 5th Text+ Plenary began around midday with the registration of participants and a group lunch, which provided an early opportunity for initial discussions and professional exchange. A total of around 130 people took part on each day of the event, using the Plenary for networking, discussion and knowledge transfer within the Text+ community and beyond.

Participants at the 5th Text+ Plenary 2026
Participants at the 5th Text+ Plenary 2026

Official opening of the 5th Plenary and introduction to the consortium’s work

The plenary was opened by Prof. Henning Lobin, Director of the Leibniz Institute for the German Language (IDS), representing the local organising team, who welcomed the guests to Mannheim Palace. In his welcome address, he emphasised the importance of sustainable research data infrastructures for the humanities and cultural studies in general, and the role of the Text+ project within the IDS in particular.

Prof. Andreas Witt, spokesperson for the Text+ consortium, then provided an overview of the project’s development, outlining its achievements to date and future challenges. To mark the conclusion of the first funding phase, he took stock of the past five years and highlighted that Text+ had established itself as a central research data infrastructure for language- and text-based research. The focus is not only on data, but also on services, standards and advisory services throughout the entire data lifecycle, consistently guided by the FAIR principles and supported by over 40 partner institutions.

Particular emphasis was placed on the integration of Text+ into the NFDI. As part of Humanities@NFDI, the consortium works closely with other NFDI consortia and is active in numerous sections and working groups. Furthermore, through its contributions to the European research infrastructures CLARIN and DARIAH, Text+ makes an important contribution to international networking.

Prof. Witt also spoke about the importance of common standards and interoperable services. With offerings such as the Text+ Data Space, federated search services, registry systems and a comprehensive service portfolio, Text+ creates the conditions for ensuring that research data remains discoverable, accessible and reusable in the long term.

Looking ahead to the second funding phase, he outlined strategic objectives such as greater integration of shared NFDI core services, closer collaboration with other consortia, and a consistent focus on the needs of the academic community. He emphasised that the success of Text+ is measured not only by its technical infrastructure, but above all by its actual use by researchers.

In conclusion, Prof. Witt highlighted the growing importance of high-quality speech and text data in the context of artificial intelligence (AI) and large language models. Against this backdrop, Text+ is well positioned to further establish itself as a key infrastructure for data- and AI-driven research in the humanities.

Task Area Presentations

The presentations from the task areas built on Prof. Witt’s remarks in his welcome address and provided insights into the results of the first funding phase.

Task Area Infrastructure & Operations

Task Area Infrastructure & Operations (IO) opened the result presentations. The task area demonstrated how the consortium’s central technical services were established and continuously developed during the first funding phase. As the technical framework for the services and offerings of the data domains Collections, Lexical Resources and Editions, IO creates the prerequisites necessary for research data to be catalogued, linked, processed, utilised, and preserved in the long term.

The initial key topics were the base infrastructure and the Search & Retrieval infrastructure. Further development of the Text+ Registry and the Federated Content Search (FCS) now enables cross-data-domain searches within the Text+ Data Space, based on both detailed metadata and the actual research content. Particular emphasis was placed on the importance of authority data for both applications and the interaction between the Registry and the FCS as a hybrid search system.

The task area then presented a selection of projects in the context of software services, data processing pipelines and central long-term archiving. These include a range of tools for Jupyter Notebook-based working environments, including a tool for a simplified data import into the TextGrid Repository and the integration of the MONAPipe NLP pipeline. Using these applications as examples, the Task Area demonstrated its active involvement in shaping the NFDI-wide base services (in this case: Jupyter4NFDI). Another focus lay on the Text+ archive as a generic fallback solution for long-term archiving, which will preserve and make available both open and restricted-access research data.

The presentation then addressed the topics of interoperability and reusability. In particular, the close collaboration with Base4NFDI was highlighted, in which the Task Area plays an active role – through, amongst other things, participation in the relevant committees, the integration of available base services and involvement in incubator projects – thereby continuously advancing the interoperability of Text+ within the NFDI. The [GND Agency Text+](<https://text-plus.org/en/themen-dokumentation/standardisierung/gnd-agentur/), launched in 2024, was also showcased; it supports researchers and projects in enriching their data using authority files. The developed workflow and the entityXML exchange format provide key components for the standardised collection and interlinking of research data. This section concluded with information on the TA’s services in the field of research data management supporting the use of standards and the sustainable reuse of research data.

The last section was devoted to the TA’s communication and community outreach. A mayor focus was put on the development of the Text+ web portal as the central point of access to the consortium’s services and information for researchers. The task area actively supports knowledge transfer through consulting, event series such as the IO Lectures and “Show and Tell”, and other exchange formats. Finally, a brief outlook on the next project phase was provided, which will centre on the further expansion of the Text+ Data Space as a continuously growing, interconnected data space that is easily accessible to researchers. This includes, amongst other things, the integration of further datasets, the improved interlinkage between datasets, and the increased use of GND identifiers and other authority data.

Further information and detailed content can be found in the presentation on PDF, pages 25–90.

Task Area Editions

The afternoon programme on the first day of the event was organised by the Task Area Editions. The focus was on the collaboration between Text+ and the Editions community. The session began with a presentation on the Task Area’s key areas of work – Community Activities, Text+ Consulting and Text+ Registry.

The session began with a joint review with the participants of the user stories and the numerous events from the first funding phase. It became clear that it is not only large-scale events such as the Text+ Plenary or conferences and association meetings – for example, the DHd annual conferences, the DARIAH Annual Event, the Cologne Medieval Studies Conference or the annual conference of the Working Group on Germanistic Editions – that strengthen collaboration within the community. In particular, smaller formats, tailored to the needs of researchers – such as FAIR February or the ‘Research Rendezvous’ – as well as numerous workshops and advisory services for and with editors that have a lasting impact.

The second focus of the presentation was on the Text+ Consulting service area. The research data management and grant application advice provided by the team of Text+ agents has now established itself as an indispensable part of the research infrastructure for edition studies and, more broadly, the humanities. A keyword-based analysis of the consultancy records yielded both expected findings – such as the high demand for Text+ support with funding applications and letters of support – and surprising insights, for example into the now routine integration of AI into editorial workflows.

A particular success has been the consistent further development of Text+ Consulting through networking with the relevant structures of NFDI4Culture, NFDI4Memory and NFDI4Objects. Together, this creates a networked advisory service for the humanities-based NFDI consortia, which seamlessly routes FDM enquiries to the appropriate advisers, regardless of subject area or specific focus.

The open catalogue of editions, as part of the Text+ Registry, was then presented as the technical centrepiece of the task area. Over the past five years, the catalogue has been further developed into a central meta-catalogue that increasingly serves as an interoperable backbone for the documentation and networking of scholarly editions.

The TA Editions session concluded with a panel discussion on the topic ‘What does the editorial community need? Perspectives in Text+ and NFDI 2.0’. Building on the insights from the preceding presentation and brief opening remarks on the question ‘What characterises good editorial infrastructure?’, Prof. Andrea Rapp (Co-Lead of the TA Editions), Dr. Hannah Busch (SCC Editions), Swantje Dogunke (SCC Editions) and Dr. Philipp Hegel (TA Editions) discussed the impact Text+ has had so far within the editorial community and the areas where there is still potential for development. The panel was moderated and hosted by Kilian Hensen (Coordinator TA Editions). Questions and suggestions from the audience were incorporated into the discussion using the arsnova engagement tool. The key discussion points and outcomes will be published in the coming weeks in a separate post on the Text+ blog.

The Task Area Editions would like to thank all the panellists and the audience for their lively, constructive and critical participation.

Further information and detailed content can be found in the presentation on PDF, pages 92–129.

Poster session

The official part of the first day of the event concluded in the late afternoon with a poster session, which began with a poster slam.

A total of 18 posters were accepted for the Text+ Plenary. The aim was for the posters to showcase a broad spectrum of topics within Text+, covering key areas ranging from metadata and NFDI collaborations, through Text+ funded collaborative projects and community work, to the topics of AI and tools:

Metadata

The posters relating to the field of metadata dealt with the harmonisation and further development of metadata standards within Text+ and the NFDI. The focus was on uniform models such as CMDI for lexical resources, as well as the results of interdisciplinary workshops on an NFDI-wide Core Metadata Profile. In addition, domain-specific requirements, such as those for literary corpora, were presented with the aim of ensuring better interoperability and reusability of research data.

  1. A Unifying CMDI-based Metadata Schema for Lexical Resources in Text+ (DOI: https://doi.org/10.5281/zenodo.20811138)

  2. From Interdisciplinary Metadata Alignment to Practice in Text+: Results and Implications of the 2nd NFDI Metadata Workshop (DOI: https://doi.org/10.5281/zenodo.20627968)

NFDI Collaborations

This part of the poster session focused on the integration of Text+ with other NFDI consortia and core services. Topics included shared infrastructures such as Humanities@NFDI and Base4NFDI, as well as interfaces with BERD and other partners. Examples illustrated cross-consortium use cases and the integration of core services into Text+.

  1. Bring Your Own Topic: Developments in the open Text+ ‘Show and Tell’ event series (DOI: https://doi.org/10.5281/zenodo.21129812)

  2. Bridges between business, economic and text data: Collaboration between BERD and Text+ within the NFDI (DOI: https://doi.org/10.5281/zenodo.20924254)

  3. Humanities@NFDI (DOI: https://doi.org/10.5281/zenodo.16835734)

  4. Integration of Basic Services in Text+ (DOI: https://doi.org/10.5281/zenodo.18999643)

  5. Restoration documents as research data – annotation, knowledge modelling and inter-consortium perspectives at the interface between NFDI4Objects and Text+ (DOI: https://doi.org/10.5281/zenodo.20732860)

  6. Synergies between Base4NFDI services and Text+ (DOI: https://doi.org/10.5281/zenodo.20812397)

Collaborative Projects (CP)

The two collaborative projects presented, which were funded by Text+, demonstrated the practical implementation of standards within Text+. The Transcription+ collaborative project focuses on the standardisation of spoken language and tool support, whilst AWV 2.0 makes historical data sets accessible again. Both projects illustrate how heterogeneous or outdated data can be made FAIR-compliant through migration and standardisation.

  1. Transcription+ ISO 24624:2016 – Transcription of spoken language: resources, documentation and a multilingual demo corpus (DOI: https://doi.org/10.5281/zenodo.20848185)

  2. Making inaccessible research data FAIR again – the AWV 2.0 collaboration project (DOI: https://doi.org/10.5281/zenodo.20828973)

Community

The posters on the topic of community presented formats for user-centred approaches and scientific communication. These include blogs, event series, usability workshops, as well as helpdesk and advisory services.

  1. AG Dissemination on tour (DOI: https://doi.org/10.5281/zenodo.21277140)

  2. Community Outreach through Scientific Blogging: the Text+ Blog (DOI: https://doi.org/10.5281/zenodo.20591639)

  3. Let’s talk to the users. Usability workshops for the user-centred further development of the FCS (DOI: https://doi.org/10.5281/zenodo.21129880)

  4. Close to research: Text+ as a federated infrastructure for RDM advice and helpdesk support (DOI: https://doi.org/10.5281/zenodo.20842616)

AI & Tools

This part of the poster session covered AI-supported methods and digital tools for research data. These included RAG systems, NLP pipelines and LLM applications in lexicography, as well as infrastructure components such as registries and repositories. Together, the posters provided examples of how technical innovation can improve access to, analysis of and re-use of text-based data.

  1. CarMa Bot: Evaluation of RAG approaches for content exploration in editions of musicians’ letters (DOI: https://doi.org/10.5281/zenodo.21114371)

  2. MONAPipe: Modular Natural Language Processing Pipeline for Digital Humanities (DOI: https://doi.org/10.5281/zenodo.8424925)

  3. The Text+ Registry: Facilitating Cross-Domain Resource Discovery (DOI: https://doi.org/10.5281/zenodo.20036812)

  4. Using large language models to generate example sentences for GermaNet (DOI: https://doi.org/10.5281/zenodo.20810678)

Evening event

Following informative presentations, lively discussions and the poster session, a relaxed evening reception featuring regional food and drinks took place in the Rectorate Courtyard of Mannheim Palace, in glorious summer weather. The event provided an opportunity for informal exchanges and in-depth discussions in a pleasant atmosphere. The majority of participants accepted the invitation to this convivial get-together.

Evening reception at the 5th Text+ Plenary on 15 June 2026
Evening reception at the 5th Text+ Plenary on 15 June 2026

Day 2 of the programme: 16 June 2026

Task Area Presentations

The morning of the second day of the event, which was open to the public, focused on the presentations from the Lexical Resources and Collections task areas:

Task Area Lexical Resources

The presentations from the Task Area Lexical Resources provided a comprehensive review of the results of the first funding phase of Text+. The Task Area brings together the key players in digital lexicography and lexical research in Germany and now provides 94 lexical resources in more than 30 languages, language levels and varieties for academic reuse. The resources are used across a wide range of disciplines – from linguistics and literary studies to the digital humanities, computational linguistics and AI research.

A key focus of the presentation was on the seven collaborative projects funded by Text+ in the ‘Lexical Resources’ area. These significantly expanded the range of resources available within Text+ and included, among other things, dictionaries and lexical data on Middle High German, Lower Sorbian, Latin, Greek-Arabic and Ancient Egyptian. At the same time, the projects provided valuable impetus for the further development of common standards and technical solutions within Text+.

Particularly impressive was the presentation of various application scenarios for the shared dictionary platform, which is based on the Lexical Content Search (LexFCS) developed in the Task Area, the curated lexical data – much of which has been newly compiled – and the Text+ Registry. A live demonstration showed how researchers can access distributed and highly heterogeneous dictionaries and word networks via a shared web interface. A second, more technically oriented demo highlighted the even broader possibilities offered by the direct use of the dictionary platform’s application programming interfaces (APIs) and their direct integration into individual research projects. This demonstrated how semantic relationships and shared identifiers can be used to navigate between different lexical resources and to link content across institutions and systems. The demonstration vividly illustrated Text+’s vision of making research data not only discoverable, but also interoperable and usable for new research questions.

In addition, the Task Area presented its work on the harmonisation of metadata, the further development of the LexFCS specification and its implementation, the Text+ Registry, and its contributions to European standardisation within the CLARIN and DARIAH frameworks. A uniform metadata profile based on CMDI (ISO 24622) will facilitate the integration and interlinking of lexical resources in future.

Another key focus was the growing importance of artificial intelligence for working with lexical resources. Initial experiences with large language models in the creation of example sentences for GermaNet were presented, along with transformer-based conversion of historical texts into modern orthography and new approaches to AI-supported search queries.

With a view to a possible second funding phase, the joint dictionary platform is to be further expanded and more closely integrated with other data domains within Text+. The aim is to evolve from a distributed resource catalogue into an integrated research platform that provides both human users and AI-based applications with seamless access to lexical knowledge.

Further information and detailed content can be found in the presentation on PDF, pages 132–191.

Task Area Collections

The presentation by the Task Area Collections was guided by an unusual theme: rather than presenting the results of the past five years as a traditional project review, the Task invited the audience on a ‘guided tour of the greatest achievements of the TA Collections over five years of Text+’. In the form of a museum tour, key developments, projects and infrastructures were presented as exhibits within a shared collection.

The starting point was the question of how Text+ could ensure that language- and text-related research data remain visible, discoverable and reusable in the long term. The Task Area’s numerous collections and data centres were presented, as well as the continuously expanding Text+ Registry, which now lists almost 1,500 resources spanning a wide variety of languages, data types and subject areas. Together with Federated Content Search, it enables cross-location access to research data from libraries, archives, repositories and academic projects.

Several examples were used to demonstrate how existing and emerging data collections are being integrated into the Text+ infrastructure. These include, amongst others, the Göttingen Register of Electronic Texts in Indian Languages (GRETIL) in the TextGrid Repository. The presentation illustrated how heterogeneous and, in some cases, difficult-to-access collections can be made FAIR-compliant in the long term through metadata enrichment, standardisation and integration into established repositories.

Another exhibit in the ‘collection’ was the Text+oh collaboration project, funded by Text+ in the Collections data domain. The presentation showcased the integration of the Oral-History.Digital portal, which provides access to thousands of audiovisual interviews from numerous archives. By integrating the metadata into the Text+ Registry and converting transcripts into standardised formats, the visibility and reusability of these important audiovisual research resources are significantly improved.

Looking beyond the research community itself, the Task Area deliberately chose to look beyond traditional boundaries. Through workshops and discussion forums, it demonstrated how Text+ collaborates with partners from industry, further education and forensics to open up new fields of application for language and text data and to address the needs of different user groups. This outreach to other sectors of society was highlighted as a key component of Text+’s community work.

A particular highlight was the presentation of the new DIN standard 19461 on derived text formats. In keeping with the exhibition concept, the standard itself was treated and interpreted as an ‘exhibit’. The speakers demonstrated how controlled transformations of copyright-protected original texts can result in research data that can be reused in a legally compliant manner. The standard was presented not merely as a technical set of rules, but as a vital infrastructure for research that is both open and legally compliant.

The presentation was rounded off with insights into new tools and AI-based services, including the MONAPipe NLP platform, entity linking methods and the Text+ chatbot. The developments showcased highlighted just how closely the work of the Task Area Collections is now linked to current AI applications, research data standards and the long-term provision of scientific resources.

Further information and detailed content can be found in the presentation on PDF, pages 193–264.

Outlook and Closing Remarks

The official programme for the second day of the event concluded with a presentation by the co-spokesperson of Text+ for the first funding phase, Prof. Philipp Wieder of the Georg-August University of Göttingen, in which he outlined the possible future development of Text+. Against the backdrop of announced funding cuts in the event of continued funding, as well as changing framework conditions, he outlined a strategic reorientation of the project. The central question addressed in his remarks was the key question of what would make Text+ indispensable to research. He emphasised that, in future, the focus should be less on the range of services offered and more on their quality and actual use in research processes. He also highlighted the need for greater integration of Text+ into academic practice. Overall, the focus of Text+ should shift from quantity to quality, from what is on offer to how it is used, and towards greater visibility within the academic community.

Prof. Wieder also explained that the Task Areas would remain the structural foundation of Text+ and would continue to cover the full breadth of text- and language-based research. He also announced that, in a second funding phase, the number of Task Areas would be expanded from four to five: in addition to TA Editions, TA Lexical Resources and TA Infrastructure/Operations, TA Collections would be divided into two: TA Library holdings as Research Data, for the structured re-use of library holdings, and TA Language Corpora and Elicited Data, for corpora and elicited data. The administrative Task Area will remain in place and will continue to be responsible for coordination, communication and networking.

Furthermore, Prof. Wieder highlighted the new opportunities arising from the development of large language models: high-quality, curated data is becoming increasingly crucial – he sees this as a particular strength of Text+. This will enable Text+ to further establish itself as a reference infrastructure for data- and AI-driven research.

In conclusion, it was emphasised that the next two years will be crucial for bringing together infrastructure, expertise and the community in such a way that Text+ becomes indispensable in day-to-day research – in collaboration with the community and tailored to its research needs.

Lunch

The public programme on the second day of the event concluded with a group lunch.

Internal Meetings

In addition to the public programme, several internal working meetings took place on the second day of the event: The coordination committees for the Task Areas Collections and Editions each held meetings. The Task Area Lexical Resources also met with its coordination committee. In addition, the FID Working Group met for an internal discussion.

Thanks

We would like to extend our sincere thanks to all participants and contributors for their active involvement, open discussions, and constructive discussions and exchanges. Special thanks go to the coordinators of the Task Areas, whose dedication and substantive contributions were essential to the success of the 5th Text+ Plenary.