Oct 05, 2026
Listen On:

Episode 50: Search Is the Interface: Semantic Search and RAG - Olivier Dobberkau (dkd / TYPO3)

--:--
--:--
--:--
--:--

Enterprise search is becoming an interface rather than a simple keyword box. Olivier Dobberkau, CEO and founder of dkd Internet Service GmbH and president of the TYPO3 Association, argues that organizations should improve content quality, metadata and retrieval before adding AI-generated answers.

In Episode 50 of Beyond the CMS, Olivier joins Chris Bryce to explain how website search is evolving from keyword matching toward semantic search, retrieval-augmented generation (RAG), intent routing and specialized search capabilities. His central message is practical: AI search cannot compensate for unclear, outdated or poorly governed content.

Answer in brief: To prepare an enterprise website for AI search, collect real user questions, study zero-result searches, repair and annotate the underlying content, strengthen the retrieval layer, and only then add generated answers with grounding, monitoring and feedback.

Search is becoming the interface

Traditional website search asks a visitor to enter keywords, ranks matching documents and returns a results list. Visitors are now increasingly familiar with systems that accept complete questions and return direct answers. That changes what people expect from enterprise websites.

Instead of forcing visitors to guess an organization’s terminology, a search experience can interpret intent and help someone find an answer, document, person, price or next step. Olivier also anticipates websites serving both people and programmatic agents, increasing the importance of explicit structure, metadata and relationships.

Why content quality comes before AI search

AI search still depends on the source material it retrieves. If content is ambiguous, duplicated, outdated or written only in internal language, the system has weak evidence from which to build an answer.

Olivier recommends reviewing what visitors actually ask, examining searches that return no useful results and checking whether website content uses the language of its audience. Clear headings, concise definitions, answer-first writing and useful metadata make content easier for people and machines to interpret.

Keyword search, semantic search and hybrid retrieval

Keyword search primarily matches terms in indexed fields such as titles and body copy. A well-configured index can rank those matches effectively, but it may struggle when a visitor’s language differs from the source document.

Semantic search represents the meaning of documents and questions so it can locate conceptually related information without requiring exact wording. These approaches are complementary. An enterprise implementation may combine lexical matching, structured fields, filters, metadata and semantic similarity according to the audience and use case.

What retrieval-augmented generation does

Retrieval-augmented generation combines search with language-model generation. A system retrieves relevant source material, supplies that context to a model and uses it to generate an answer.

  1. Index the source material: Store documents and useful metadata in a retrieval system.
  2. Interpret the question: Transform the request into one or more lexical or semantic queries.
  3. Retrieve relevant context: Select potentially useful documents or passages.
  4. Generate an answer: Ask a language model to answer using that retrieved evidence.
  5. Show provenance: Link the answer to authoritative sources so the user can verify it.

RAG can improve grounding, but it does not guarantee accuracy. Reliability still depends on source authority, retrieval quality, chunking, ranking and appropriate review.

Why chunking can remove context

Long documents are commonly divided into smaller passages for retrieval. Those passages are easier to index and fit into a model’s context window, but a chunk can become disconnected from the section, document or policy that gives it meaning.

Teams should test whether retrieved passages preserve enough context to support the resulting answer. Document structure, metadata, parent-child relationships and source citations can all help.

Why a chatbot is not automatically better search

A chatbot is an interface, not a search strategy. A general bot may try to answer every type of question while handling none of them particularly well. Olivier describes a more deliberate model: identify distinct user intentions and route each to a specialized capability or authoritative data source.

A request for a price may need product data. A request for a record may need an authenticated system. A question about expertise may need a people directory. A contact request may need a location or service database. The interface should follow the job the visitor is trying to complete.

A practical implementation sequence

  1. Collect representative questions from customers, employees and search logs.
  2. Review zero-result and low-satisfaction searches.
  3. Repair unclear, duplicated and outdated source content.
  4. Improve metadata, ownership and content relationships.
  5. Test retrieval before adding generated answers.
  6. Add specialized capabilities for high-value intents.
  7. Log outcomes and collect explicit user feedback.
  8. Review failed answers and improve the content and retrieval system continuously.

What you will learn

  • Why search is becoming an interface rather than only a search box
  • How keyword search differs from semantic search and RAG
  • How Apache Solr supports search in TYPO3 projects
  • Why content quality and metadata determine AI-search quality
  • Why generic chatbots often fail website visitors
  • How specialized search capabilities can route different user intentions
  • How embeddings, chunks, retrieval and grounding fit into RAG
  • How zero-result analysis and feedback improve enterprise search

Chapters

  • 0:00 Welcome to Beyond the CMS Episode 50
  • 2:06 Olivier Dobberkau, dkd and the TYPO3 ecosystem
  • 5:32 Apache Solr search for TYPO3
  • 7:23 Why search is becoming the interface
  • 11:21 From keywords to conversational and agentic search
  • 13:47 Content quality, grounding and search quality
  • 19:46 Why chatbots are not always the right experience
  • 22:37 Specialized search capabilities and intent routing
  • 24:02 Question-based website search in practice
  • 26:19 Repairing content before adding RAG
  • 28:33 Retrieval, metadata, feedback and supervision
  • 31:25 RAG explained: indexing, embeddings and retrieval
  • 34:16 Chunking, context and grounded answers
  • 36:39 Open source, access and the TYPO3 community

Questions answered

What is the difference between keyword search and semantic search? Keyword search primarily matches terms in indexed fields. Semantic search represents the meaning of documents and questions so it can find conceptually related material even when the wording differs. Enterprise systems can combine both.

What is retrieval-augmented generation? RAG retrieves relevant source material and provides it to a language model before the model generates an answer. The source context can improve grounding, but answer quality still depends on retrieval and source quality.

Why does content quality matter for AI search? Unclear, duplicated or outdated content makes authoritative evidence harder to retrieve and increases the risk of incomplete or misleading answers.

Should an enterprise website start with a chatbot? Not necessarily. A better starting point is to identify real questions and improve retrieval, then choose an interface appropriate to the use case.

How can teams improve RAG results? Improve source content and metadata, evaluate chunking and ranking against real questions, preserve source context and monitor failures and user feedback.

About Olivier Dobberkau

Olivier Dobberkau is CEO and founder of dkd Internet Service GmbH and president of the TYPO3 Association. His work spans open-source CMS development, TYPO3 search, Apache Solr and intelligent content discovery. Connect with Olivier on LinkedIn.

Watch and listen

Watch on YouTube | Listen on Transistor

Related from Dotfusion

Dotfusion helps enterprise and nonprofit teams modernize complex digital platforms and build the content operations needed for trustworthy search and AI experiences. Explore our headless CMS services.