Implementing Advanced RAG with Weaviate

If you feed huge raw PDFs to a Large Language Model and just hope it works, you’ll quickly run into hallucinations, slow response times, and high token costs. For reliable, production-level results, you need a strong retrieval setup. Learning to build RAG with Weaviate is a valuable skill right now. In this article, I’ll show you how to set up advanced RAG using Weaviate.

Understanding Advanced RAG and Weaviate

Before we start coding, let’s break down how it works. A basic RAG pipeline splits a document into smaller parts (chunking), turns those parts into numbers (embedding), and finds the closest matches when someone asks a question (vector search).

Advanced RAG adds Hybrid Search to the process.

Semantic vector search is great at understanding what a query means, but it often misses exact keyword matches, like a specific product ID or a rare acronym. On the other hand, keyword search (BM25) finds exact matches but doesn’t understand the deeper meaning.

Weaviate handles this by running both searches at the same time and combining the results with a setting called Alpha (α), which you can adjust.

The formula Weaviate uses under the hood looks like this:

Score = (1 - α) · BM25_Score + α · Vector_Score

This means:

  • When α = 0, you are doing a pure keyword search.
  • When α = 1, you are doing a pure vector search.
  • When α = 0.5, you balance both methods.

Many of the RAG concepts behind systems like this are covered step-by-step in my book: Hands-On GenAI, LLMs & AI Agents.

Advanced RAG with Weaviate

We’ll run Weaviate locally using Docker and use Ollama for embeddings. This setup has no API costs, and your data stays on your own machine.

Step 1: The Local Stack Setup

First, ensure you have Ollama installed and pull a lightweight embedding model:

ollama run nomic-embed-text

Next, create a docker-compose.yml file and run Weaviate locally:

services:
  weaviate:
    image: cr.weaviate.io/semitechnologies/weaviate:1.33.1

    ports:
      - "8080:8080"
      - "50051:50051"

    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: "true"
      PERSISTENCE_DATA_PATH: "/var/lib/weaviate"
      DEFAULT_VECTORIZER_MODULE: "text2vec-ollama"
      ENABLE_MODULES: "text2vec-ollama"
      CLUSTER_HOSTNAME: "node1"

Start Weaviate:

docker compose up -d

Next, install the modern v4 Python client:

pip install -U weaviate-client

Step 2: Connecting and Designing the Schema

The v4 client is very Python-friendly and uses strict typing to help prevent errors later on:

import weaviate
import weaviate.classes.config as wvcfg

# Connect to the local Weaviate instance
client = weaviate.connect_to_local()

# Create a collection (schema) with Ollama vectorization
if client.collections.exists("EngineeringDocs"):
    client.collections.delete("EngineeringDocs")

collection = client.collections.create(
    name="EngineeringDocs",
    vectorizer_config=wvcfg.Configure.Vectorizer.text2vec_ollama(
        api_endpoint="http://host.docker.internal:11434",
        model="nomic-embed-text"
    ),
    properties=[
        wvcfg.Property(name="content", data_type=wvcfg.DataType.TEXT),
        wvcfg.Property(name="source", data_type=wvcfg.DataType.TEXT),
    ]
)

print("Collection created successfully.")

Step 3: Chunking and Ingestion

With large documents, you can’t embed a whole book as one vector. You have to split it into smaller, manageable chunks.

A simple sliding-window chunking strategy looks like this:

def chunk_text(text, chunk_size=500, overlap=50):
    chunks = []

    for i in range(0, len(text), chunk_size - overlap):
        chunks.append(text[i:i + chunk_size])

    return chunks

Now let’s create a realistic technical document:

raw_document = """
Engineering Platform Documentation

System Overview
---------------
The platform serves approximately 50,000 requests per minute and consists of
microservices deployed on Kubernetes.

Caching Layer
-------------
Redis is used as the primary caching layer to reduce database load.
All API responses are cached for 10 minutes.

If Redis becomes unavailable, requests automatically fall back to PostgreSQL.

Redis timeouts are handled using exponential backoff retries.
The retry sequence is 1 second, 2 seconds, and 4 seconds.
After three failed retries, the request bypasses Redis entirely.

Database Layer
--------------
PostgreSQL is the primary transactional database.
Read replicas are used for analytics workloads.

Authentication
--------------
Authentication is implemented using JWT tokens.
Access tokens expire after 15 minutes.
Refresh tokens expire after 30 days.

Monitoring
----------
Prometheus collects system metrics.
Grafana is used for dashboards and alerting.

Logging
-------
Application logs are stored in Elasticsearch.
Logs are retained for 90 days.

Background Processing
---------------------
Apache Airflow orchestrates ETL pipelines.
Critical workflows run every hour.

Machine Learning Platform
-------------------------
Machine learning models are deployed using FastAPI services.
Feature data is stored in Feast.
Model predictions are logged for monitoring and retraining.

Incident Management
-------------------
If API latency exceeds 500 milliseconds for more than 5 minutes,
an alert is sent to the on-call engineer.

Security
--------
Sensitive secrets are stored in HashiCorp Vault.
Database connections require TLS encryption.

Disaster Recovery
-----------------
Database backups are taken every 6 hours.
Recovery Point Objective (RPO) is 15 minutes.
Recovery Time Objective (RTO) is 1 hour.
"""

Chunk the document:

document_chunks = chunk_text(raw_document)

Ingest the chunks into Weaviate using dynamic batching:

with collection.batch.dynamic() as batch:
    for chunk in document_chunks:
        batch.add_object({
            "content": chunk,
            "source": "architecture_v2.md"
        })

print("Data vectorized and stored.")

At this stage, Weaviate automatically:

  1. Sends the text to Ollama
  2. Generates embeddings using nomic-embed-text
  3. Stores vectors
  4. Stores the original text

Step 4: The Hybrid Search Query

Now let’s retrieve information. We’ll query the database with hybrid search and set the Alpha parameter to balance meaning and exact matches:

response = collection.query.hybrid(
    query="How does the application continue serving users when the cache server fails?",
    alpha=0.5,
    limit=3
)

Here’s how you can display the retrieved chunks:

for i, obj in enumerate(response.objects, start=1):

    print(f"\nResult {i}")
    print("=" * 50)

    print(obj.properties["content"])

client.close()
Result 1
==================================================

Engineering Platform Documentation

System Overview
---------------
The platform serves approximately 50,000 requests per minute and consists of
microservices deployed on Kubernetes.

Caching Layer
-------------
Redis is used as the primary caching layer to reduce database load.
All API responses are cached for 10 minutes.

If Redis becomes unavailable, requests automatically fall back to PostgreSQL.

Redis timeouts are handled using exponential backoff retries.
The retry sequence is 1 second,

Result 2
==================================================
ometheus collects system metrics.
Grafana is used for dashboards and alerting.

Logging
-------
Application logs are stored in Elasticsearch.
Logs are retained for 90 days.

Background Processing
---------------------
Apache Airflow orchestrates ETL pipelines.
Critical workflows run every hour.

Machine Learning Platform
-------------------------
Machine learning models are deployed using FastAPI services.
Feature data is stored in Feast.
Model predictions are logged for monitoring and retrainin

Result 3
==================================================
backoff retries.
The retry sequence is 1 second, 2 seconds, and 4 seconds.
After three failed retries, the request bypasses Redis entirely.

Database Layer
--------------
PostgreSQL is the primary transactional database.
Read replicas are used for analytics workloads.

Authentication
--------------
Authentication is implemented using JWT tokens.
Access tokens expire after 15 minutes.
Refresh tokens expire after 30 days.

Monitoring
----------
Prometheus collects system metrics.
Grafana is used

Summary

Moving from a junior ML learner to a senior AI engineer means focusing less on models and more on data architecture. Anyone can use an LLM API, but real engineering is about building systems that quickly find the right context, control costs, and scale safely.

Tools like Weaviate help you learn to organize and work with unstructured data.

I hope you found this article on advanced RAG with Weaviate helpful.

For more AI and machine learning tips, follow me on Instagram. My book, Hands-On GenAI, LLMs & AI Agents, can also help you grow your AI career.

Aman Kharwal
Aman Kharwal

AI/ML Engineer | Published Author. My aim is to decode data science for the real world in the most simple words.

Articles: 2228

Leave a Reply

Discover more from AmanXai by Aman Kharwal

Subscribe now to keep reading and get access to the full archive.

Continue reading