Building a Faster (Rust-Based) Python Driver for ScyllaDB
Binding Rust to Python with PyO3, zero-copy deserialization with yoke Since Python is one of the most commonly used languages for interacting with a database, having a robust ScyllaDB Python driver is essential. The first ScyllaDB Python driver was created ages ago, back when ScyllaDB was in the 2.x stage, and it’s been used ever since. As part of the ZPP Project at the University of Warsaw… SourceMCP Gateway RBAC tutorial: Add Role-Based Access Control to your AI agent with Auth0 and Apache Cassandra
This is a follow-up to “MCP Gateway tutorial: Connect an AI Agent to Apache Kafka and HTTP Server Backends,” extends the original support chatbot with Auth0-based login, role-based access control (RBAC), and an Apache Cassandra backend.
In this MCP Gateway RBAC tutorial, you will configure MCP Gateway to enforce role-based access control for an AI agent. The walkthrough adds Auth0 authentication, maps user roles to MCP Gateway Access Control Lists, connects an Apache Cassandra backend, and shows how each persona receives only the tools their role is allowed to use.
The goal is to support least-privilege access for AI agent tools at the gateway layer: Auth0 authenticates the user, MCP Gateway maps the user’s role to an Access Control List, and the agent receives only the tools that role is permitted to call.
Get the code: MCP-Gateway/shoe-store-support-chatbot-rbac
Recap: where we left offThe previous article built a support chatbot for a shoe
e-commerce company, backed by one MCP (Model Context Protocol)
Gateway Virtual Server (supportchatbot) with two
backends: ordersapi (HTTP) providing get_orders, and
supporttickets (Apache Kafka) providing
submit_request. That version had a single, shared set
of tools for every user.
This article adds Auth0 login and three distinct roles
(support-agent, merchandiser, and
readonly-auditor), plus a third backend type,
Cassandra, so each persona gets exactly the access their job calls
for, configured entirely at the MCP Gateway.
From the previous tutorial, you should already have:
- An MCP Gateway Virtual Server
(
supportchatbot). - The ordersapi (HTTP) backend with a
get_orderstool. - The supporttickets (Kafka) backend with a
submit_requesttool.
Everything involving Auth0, OAuth, roles, and Access Control Lists (ACLs) is new in this article.
Tooling: Python 3.11+, the uv package manager, and an AWS account with Amazon Bedrock model access enabled for the model you’re using, in the region you’re calling (e.g. us-east-1).
The use case: a product catalog with tiered cost visibility-
support-agent(customer-facing staff) gets product info (name, price, stock) to help shoppers, plus order lookups and ticket filing. -
merchandiser(buying/purchasing staff) gets supplier cost and margin data to negotiate pricing, commercially sensitive information relevant to their role, distinct from customer-service tooling. readonly-auditorgets full read visibility across both domains for review purposes, with action (filing a ticket) reserved for support-agent.
Role-Based Access Control means a user’s identity maps to a fixed role, and every access decision is checked against an explicit allow-list. In MCP Gateway, that role determines which tools an AI agent can discover and call, so each persona gets exactly the access their job requires: no more and no less.
This matters especially for an LLM-backed agent, for three reasons:
- It keeps every session’s capabilities predictable, no matter
how the conversation goes. The chatbot handles free-text customer
input, and however that conversation unfolds, a support-agent
session simply can’t reach
catalogdata_get_supplier_cost: it’s not part of that role’s tool set, so there’s nothing to accidentally or deliberately reach for. The boundary is a property of the session, not a judgment call the model has to get right in the moment. - It lets you express two independent things: what a role can
see, and what it can do. merchandiser can see cost data but can’t
file tickets;
readonly-auditorcan see everything but can’t act on any of it. Splitting visibility from action this way maps naturally onto how organizations already assign responsibility. - It moves enforcement to a layer that’s easy to reason about. Rather than relying on prompt instructions like “don’t answer cost questions for this role,” which work well in the common case but depend on the model’s judgment every time, the MCP Gateway checks the ACL before a tool call is honored, giving you one clear, auditable place where access is decided.

Figure 1. Routing role-based AI agent requests to HTTP, Kafka, and Cassandra backends through the MCP Gateway
Step 1: Provision the Cassandra clusterIn the Instaclustr console, Create Cluster → Apache Cassandra, choosing a data center/region and node size appropriate for a demo.
Once the cluster reaches Running, note its Data Center ID from the Details page for Step 3.
Use the console’s Connection Info tab and default superuser credentials to connect via cqlsh for the next step.
Step 2: Set up the Cassandra schema, seed data, and a least-privilege roleCREATE KEYSPACE IF NOT EXISTS catalog WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 3}; CREATE TABLE catalog.products ( productid text PRIMARY KEY, name text, description text, category text, listprice decimal, stockstatus text ); CREATE TABLE catalog.supplier_costs ( productid text PRIMARY KEY, supplierid text, unitcost decimal, marginpercent decimal, suppliercontact text ); CREATE ROLE mcpgateway_catalog WITH PASSWORD = '<password>' AND LOGIN = true; GRANT SELECT ON catalog.products TO mcpgateway_catalog; GRANT SELECT ON catalog.supplier_costs TO mcpgateway_catalog; INSERT INTO catalog.products (productid, name, description, category, listprice, stockstatus) VALUES ('SHOE-001', 'Midnight Runner', 'Black low-top running shoe', 'running', 129.99, 'in_stock'); INSERT INTO catalog.products (productid, name, description, category, listprice, stockstatus) VALUES ('SHOE-002', 'Trailblazer Boot', 'Brown waterproof hiking boot', 'hiking', 189.99, 'backordered'); INSERT INTO catalog.supplier_costs (productid, supplierid, unitcost, marginpercent, suppliercontact) VALUES ('SHOE-001', 'SUP-88', 42.00, 67.7, 'orders@northstarfootwear.com'); INSERT INTO catalog.supplier_costs (productid, supplierid, unitcost, marginpercent, suppliercontact) VALUES ('SHOE-002', 'SUP-14', 71.50, 62.4, 'sales@trailgearco.com');
Connect with the credentials and endpoint from your cluster’s
Connection Info tab before running this (see Step 1).
mcpgateway_catalog is read-only and scoped to just
these two tables: a second, defense-in-depth layer of access
control underneath the MCP Gateway’s own ACLs.
Notes:
- Identifiers are written lowercase throughout, matching how Cassandra represents unquoted identifiers internally: table/column names and named bind markers alike. Writing them lowercase from the start keeps the CQL here consistent with the tool configuration in Step 3.
- Remember to replace <password> in the
CREATE ROLE mcpgateway_catalogstatement with a real, strong password before running this; it’s a placeholder, not a literal value. Keep it somewhere secure, since you’ll need to re-enter it in Step 3 when you create thecatalogdatabackend on the MCP Gateway (that’s where these credentials actually get used).
On the supportchatbot MCP Virtual Server → Create
Backend → name catalogdata, type Cassandra, using the
Data Center ID from Step 1 and the mcpgateway_catalog
credentials from Step 2.

Figure 2. The catalogdata Cassandra backend,
showing its Cassandra Data Center ID, Cassandra Username
(mcpgateway_catalog), and its two configured
tools.
Open the new catalogdata backend and, under Tools,
click Create Tool to create the below two tools.
get_product_details:
SELECT name, description, category, listprice, stockstatus FROM catalog.products WHERE productid = :productid
- Parameter:
productid(string) - Result fields:
name(string),description(string),category(string),listprice(string),stockstatus(string) - Description: “Look up a product’s customer-facing details by product ID, including its name, description, category, list price, and current stock status.”
get_supplier_cost:
SELECT supplierid, unitcost, marginpercent, suppliercontact FROM catalog.supplier_costs WHERE productid = :productid
- Parameter:
productid(string) - Result fields:
supplierid(string),unitcost(string),marginpercent(string),suppliercontact(string) - Description: “Look up internal supplier cost and margin data for a product by product ID. This is sensitive commercial data; only use it when explicitly asked, and only if you have access to this tool.”

Figure 3. The get_product_details tool definition,
showing its CQL query, productid parameter, and result
fields.
Notes:
- Declare Result Field Type as string for any Cassandra decimal column.
- Give every tool a real, descriptive description. Beyond helping the model choose the right tool, AWS Bedrock’s Converse API expects every tool in the list to have one, so it’s worth treating as a required field alongside the CQL and parameters.
- Apache Cassandra forces unquoted bind variable names to
lowercase:
:productIdis interpreted internally as:productid. That’s why every identifier in this tutorial (table/column names and bind markers alike) is written lowercase from the start, keeping the CQL here consistent with the tool configuration above. If you ever need a case-sensitive parameter name, quote it in the query instead, e.g.:"productId".
4a. Application and Post-Login Action: Follow the Auth0 example configuration guide to:
- Create a Regular Web Application and configure its Allowed Callback URIs / Web Origins.
- Create the Post-Login Action that adds a roles claim when the
mcp_rolesscope is requested, add theMCP_GATEWAY_CLIENT_IDsecret, deploy it, and apply it on the post-login trigger. - Note the audience (API Identifier), issuer, and JWKS URI as described there; needed in Step 5.
4b. Create the three Auth0 Roles: User Management → Roles → Create Role, once each for:
support-agentmerchandiserreadonly-auditor
4c. Create three test users, one per role: User Management → Users → Create User (or reuse existing test users), then on each user’s Roles tab, Assign Roles with the matching role from 4b.
Note: Keep role names consistent,
case-for-case, across three places: the Auth0 Role, the
user_roles claim value it produces, and the MCP
Gateway ACL’s Role name field in Step 6. Matching these exactly is
what makes each persona’s tool list show up correctly.
On the supportchatbot MCP Virtual Server’s OAuth
configuration section (refer to the documentation for
more details), enter the values noted in Step 4a:
- Issuer:
https://<domain>/ - JWKS
URI:
https://<domain>/.well-known/jwks.json - Audience: your API Identifier
- Scopes Supported:
mcp_roles - Roles Claim Name:
https://mcp_gateway/user_roles
Save/apply.

Figure 4. The supportchatbot MCP Virtual Server’s
OAuth configuration with all three backends.
On the MCP Gateway’s ACL screen, you pick each tool by its bare name; the Backend column disambiguates which tool you mean when the same name could theoretically exist on more than one backend:
support-agent
Tool Backendget_orders ordersapi
submit_request supporttickets
get_product_details catalogdata
readonly-auditor
Tool Backendget_orders ordersapi
get_product_details catalogdata
get_supplier_cost catalogdata
merchandiser
Tool Backendget_product_details
catalogdata get_supplier_cost
catalogdata

Figure 5. The support-agent Access Control List, showing its explicitly allowed tools.
Each role has a distinct tool set tailored to its purpose: merchandiser focuses entirely on catalog data, while readonly-auditor sees everything the other two combined can see, with action reserved for support-agent.
Once exposed through the supportchatbot MCP Virtual
Server, the agent sees each tool under a backend-prefixed name
(e.g. catalogdata_get_product_details) to disambiguate
tools coming from different backends; that’s the name that shows up
in the app’s “Tools discovered” sidebar later, and in Bedrock’s
tool-calling requests, even though you select it by its bare name
here.

Figure 6. The completed supportchatbot MCP Virtual
Server, showing its OAuth configuration, all three backends, and
the three role-based Access Control Lists.
None of this required touching the chatbot’s code, because it
was built to discover tools dynamically rather than hardcode
them: mcp_manager.py‘s
connect() attaches the user’s access token as a Bearer
header once, at connection time, and every session it opens rides
on that same authenticated connection for the rest of the login;
get_tool_specs() then simply asks the MCP Gateway
which tools this identity can see and turns whatever comes back
into Bedrock’s toolConfig.tools format, with the MCP
Gateway itself deciding (via the caller’s role and its Access
Control List) which tools are included; and app.py‘s
sidebar just renders that same list, grouping tool names by server
with nothing hardcoded. So, when catalogdata‘s two
tools were added at the MCP Gateway, they simply started showing up
for the roles allowed to see them, with zero code changes anywhere
in this repo.
The full code for this tutorial is available at MCP-Gateway/shoe-store-support-chatbot-rbac in
the instaclustr/code-samples repo.
cp config.example.json config.json # fill in your Auth0 tenant + MCP Gateway URL uv sync export AWS_ACCESS_KEY_ID="..." export AWS_SECRET_ACCESS_KEY="..." export AWS_DEFAULT_REGION="us-east-1" uv run streamlit run src/rbac_shoe_bot/app.py
This opens the app at http://localhost:8501. Compared to the
shared-access version from the previous tutorial,
config.json now includes a populated auth0 section,
and the app opens on a Log in with
Auth0 screen.
Click Log in with Auth0 and sign in as one of your test users. To try a different persona, click Log out and log in again; each login prompts for credentials fresh, so you can move between test users freely.

Figure 7. A readonly-auditor session in the
chatbot, showing its discovered tools in the sidebar and a
successful get_supplier_cost lookup for
SHOE-002.
Notes:
- Enable model access for
bedrock.model_idin the Bedrock console for your region ahead of time; it’s a quick, one-time opt-in per model/region, and worth doing before your first run. - The sidebar’s Tools discovered list is the fastest way to
confirm your ACL setup is working as intended; a
merchandisersession should show only the twocatalogdata_*tools.
Demo script
Using the two seeded products:
support-agent: “What’s the price and stock status of product SHOE-001?” → succeeds. “What do we pay our supplier for it?” → no tool available for this role.merchandiser: “What’s our cost and margin on SHOE-001?” → succeeds. “Show me my recent orders.” → no tool available for this role.readonly-auditor: sees both cost and orders, with ticket filing reserved forsupport-agent.
Note: These prompts reference product IDs
directly rather than descriptions like “black shoes,” since
get_product_details looks up by exact
productid. Referencing known IDs keeps the demo
focused on RBAC rather than search.
Authentication, three distinct roles, and a whole new backend technology were all added here through configuration: Auth0 setup, MCP Gateway OAuth config, ACLs, and the Cassandra backend/tools, without changing the chatbot’s own code. That’s the strength of centralizing access control at the MCP Gateway layer: the same access rules apply consistently, however a conversation unfolds, because the boundary lives outside the agent’s own reasoning.
Try it yourselfIf you want to add role-based access control to your own AI agent, you can spin up MCP Gateway and an Apache Cassandra cluster today: start a free Instaclustr trial, no credit card required. From there:
- Follow the MCP Gateway documentation to provision a Virtual Server and connect your first backend.
- Use the Auth0 example configuration guide to wire up your own identity provider.
- Clone MCP-Gateway/shoe-store-support-chatbot-rbac and swap in your own tools, roles, and backends: the same three steps (Auth0 roles claim, Gateway OAuth config, per-role ACLs) apply no matter what your agent actually does.
If you’d rather start from the simpler, single-role version, the previous tutorial walks through connecting an agent to Kafka and HTTP backends without the auth layer.
The post MCP Gateway RBAC tutorial: Add Role-Based Access Control to your AI agent with Auth0 and Apache Cassandra appeared first on Instaclustr.
Vector Search in Apache Cassandra for RAG Retrieval
Semantic retrieval using one Cassandra table to understand how it works, where it shines, and what to watch for.
Vector search in Apache Cassandra ® lets teams store embeddings, index them with Storage-Attached Indexing, and run approximate nearest neighbor queries directly in CQL without managing a separate vector database. For RAG retrieval, that means Cassandra can act as both the operational database and the semantic search layer when the retrieval problem is primarily meaning-based.
Search used to require the person asking to already know the answer’s vocabulary. A technician precisely asks, “Recall 24V-113: tailgate latch”; a truck owner asks, “Is there a recall on the tailgate?” Embeddings close that gap by mapping meaning into geometry, so similar ideas land near each other in vector space regardless of the words used. That geometry is the retrieval layer underneath most AI applications. Using retrieval-augmented generation (RAG), a pattern for answering questions from supplied data instead of the model training set alone, the chatbot can be instructed to only answer from the documents vector search hands it. As a result, the retrieval quality sets the ceiling on the answer.
The standard approach utilizes a dedicated vector store alongside the operational database. While Cassandra is not the first operational database to keep embeddings next to the data they describe, version 5.0 made it a native part of the data model. When the use case fits, querying embeddings on the same infrastructure that houses the data can cut both complexity and cost: one database to operate instead of two, and no application code responsible for keeping them in sync.
This is Part 1 of 2. Part 1 demonstrates semantic retrieval against a realistic knowledge base containing vehicle recalls, tow ratings, and warranty terms using Cassandra. Then, we show where Approximate Nearest Neighbor (ANN) alone is not the complete answer, which sets up a different approach, hybrid RAG, in Part 2.
What is vector search on Apache Cassandra?Vector search on Apache Cassandra lets you store, index, and
query high-dimensional embeddings natively in the database using
the vector<float, N> column type and
Storage-Attached Indexing. Introduced in Cassandra 5.0, it
enables ANN search through an ORDER BY … ANN
OF clause in CQL, with no separate vector store.
At write time, title + body is embedded and stored in a vector
column on the same row as the text. At query time the question is
embedded with that same model, and Cassandra is asked for the
nearest stored vectors. Three pieces of CQL do the whole job:
declare the column, attach an SAI index so ANN is possible, and
query with ORDER BY … ANN OF plus an optional
similarity score in the select list.
CREATE TABLE docs ( id text PRIMARY KEY, title text, body text, category text, model text, model_year int, embedding vector<float, 384> ); CREATE INDEX ON docs (embedding) USING 'sai' WITH OPTIONS = {'similarity_function': 'COSINE'}; SELECT id, title, category, model, model_year, similarity_cosine(embedding, ?) AS score FROM docs ORDER BY embedding ANN OF ? LIMIT 5;
Each piece of that, in turn:
vector<float, N>is a native CQL column type holding an embedding of N dimensions, sitting on the row it describes rather than in another system. It’s 384 here becauseall-MiniLM-L6-v2(which is used for the demo application) produces 384-dimensional vectors; document and query vectors must always use the same model and dimension.- SAI (Storage-Attached Indexing) attaches a searchable index to the table as rows are written. For vector columns it plugs in JVector, an ANN engine that walks a graph of nearby embeddings for fast top-k retrieval.
-
similarity_functionacceptsCOSINE(the default),DOT_PRODUCT, orEUCLIDEAN. Changing it later means dropping and recreating the index. ORDER BYembeddingANN OF ?returns approximate nearest neighbors ordered by similarity. This is fast, but not a guarantee of the true global nearest neighbor.similarity_cosine(embedding, ?)is optional and exists to make the ranking legible. It returns a 0-to-1 value rather than raw cosine, so 0.5 means a document unrelated and a score of 0.75 is not “three-quarters relevant.”
Writes don’t need special handling because the embedding binds
as one more parameter on an ordinary
prepared INSERT alongside the document, and
SAI indexes it as the row is written, under the replication factor
and consistency level you already use. In addition, vectors impose
no special replication requirements. The demo application’s
keyspace is SimpleStrategy with RF 1 only because it’s a
single-node Docker cluster.
Throughout this series, you own a fictional 2024 Summit 1500 and you’re asking a support bot a question. The knowledge base is fictional but realistic: 21 documents covering 12 recalls, tow ratings, payload figures, maintenance intervals, and warranty terms for Summit 1500 and 2500 pickups.
The recall documents matter more than their count suggests. They share one sentence skeleton, mimicking how a real recall notice would read, making them a clean test of something specific: whether an embedding can resolve an identifier when the surrounding prose is nearly identical.
Everything in this series runs on Apache Cassandra 5.0.9 in
Docker, with embeddings generated locally from the all-MiniLM-L6-v2, an
embedding model, so there’s no API key to supply.
Companion app: cassandra-vector-demo.
Follow the instructions in the README to reproduce the output
below. Running the application isn’t required to understand the
concepts in this series.
Throughout this demo, you never directly query Cassandra. Rather, you’ve asked a question, expect an answer back, and never see the retrieved document’s similarity scores. Vector search is the middle step that chooses which rows the language model is allowed to read:
- The application takes the question as text.
- The same embedding model used at insert time turns that question into a vector.
- Cassandra returns a small list of nearest documents
(
ORDER BYembeddingANN OF … LIMIT k). - The application copies those titles and bodies into a prompt, next to the original question. That block is the context.
- The language model writes an answer using (in a well-behaved RAG setup) only that context. It doesn’t scan the knowledge base, only the handful of chunks retrieval picked.
We run three Summit owner questions against the table above and
read what Cassandra ranks. The demo doesn’t call an LLM after
retrieving the results; the point is whether they are trustworthy
before an LLM ever sees it. The first question is a paraphrase,
which is what embeddings exist to solve. The other two are the
limits you plan around: an exact identifier, and a question too
vague to have one answer. Both are properties of similarity search
rather than of Cassandra; the same model against Pinecone,
pgvector, or OpenSearch kNN should produce the same or similar
ordering and both are fixed with a query change rather than a
database change. The output below is generated
from src/demo.py against Cassandra 5.0.9
with all 21 documents seeded.
“Is there a recall on the tailgate?”
python src/demo.py "Is there a recall on the tailgate?"QUERY: Is there a recall on the tailgate? 1. [recalls] recall-24v-113 [1500 2024] Recall 24V-113 tailgate latch score=0.7767 2. [recalls] recall-23v-011 [1500 2023] Recall 23V-011 instrument cluster score=0.7032 3. [recalls] recall-23v-088 [1500 2023] Recall 23V-088 backup camera score=0.6935 4. [recalls] recall-23v-145 [2500 2023] Recall 23V-145 brake booster score=0.6934 5. [recalls] recall-22v-078 [2500 2022] Recall 22V-078 trailer brake module score=0.6808
Vector search ranks Recall 24V-113 (tailgate latch) first, and the margin is wide: 0.7767 against 0.7032 for the runner-up. Nobody wrote the word “tailgate” into the question expecting a latch recall, and the model bridged it anyway, against 11 other recall documents written in nearly identical prose. Everyday paraphrase is exactly what embeddings are for, and one CQL query against one table is the whole implementation.
Outcome 2: An exact recall ID lands nowhere near the top“What is recall 24V-330?”
Now you have the recall number.
python src/demo.py "What is recall 24V-330?" QUERY: What is recall 24V-330? 1. [recalls] recall-23v-011 [1500 2023] Recall 23V-011 instrument cluster score=0.7871 2. [recalls] recall-24v-207 [1500 2024] Recall 24V-207 fuel pump relay score=0.7701 3. [recalls] recall-23v-088 [1500 2023] Recall 23V-088 backup camera score=0.7623 4. [recalls] recall-24v-113 [1500 2024] Recall 24V-113 tailgate latch score=0.7560 5. [recalls] recall-24v-155 [2500 2024] Recall 24V-155 glow plug controller score=0.7514
The document for 24V-330 sits at rank 9, scoring 0.7332, below eight notices that never mention it. The recall notices share one sentence skeleton, so the ID is a handful of characters inside an otherwise near-identical embedding, and ANN ranks by overall closeness. Cosine distance has no notion of an exact token match, which is why the same model produces this ranking in any vector store, and why the fix is a lexical path beside ANN rather than a better embedding model. Part 2 of the series works to address this limitation.
Outcome 3: A question retrieval cannot answer as asked“How much can the Summit 1500 tow?”
python src/demo.py "How much can the Summit 1500 tow?" QUERY: How much can the Summit 1500 tow? 1. [towing] summit-1500-2024-tow [1500 2024] Tow ratings by configuration, 2024 model year score=0.8990 2. [towing] summit-1500-2023-tow [1500 2023] Tow ratings by configuration, 2023 model year score=0.8853 3. [towing] summit-2500-2024-tow [2500 2024] 2024 Summit 2500 heavy duty towing capacity score=0.8791 4. [specs] summit-1500-2024-payload [1500 2024] 2024 Summit 1500 payload capacity score=0.8276 5. [towing] hitch-classes Hitch receiver classes and weight distribution score=0.7550
The ranking looks reasonable until you read the bodies. The top four sit within 0.08 of each other and answer four different questions: 11,300 lb for the 2024 1500, 9,100 lb for the 2023 1500, 17,600 lb for the 2500, and 1,750 lb of payload, which isn’t a tow rating at all.
Retrieval did its job since all four are genuinely relevant to
what was asked. The question just never supplied a model year, so
nothing in the index can pick a winner. The fix is
a WHERE model_year = ? alongside ANN, which
is exactly why those columns are on the table, and will be utilized
in Part 2.
Vector search is enough when the question is about meaning
and an incorrect response is recoverable: everyday
paraphrase like “is there a recall on the tailgate,” similar
tickets, or similar products already sitting in Cassandra. It isn’t
enough when an identifier, a model year, or a customer record has
to be exact, and no similarity threshold will separate those two
cases for you. Those need a lexical path or
a WHERE clause next to ANN, which is the
work of Part 2 using additional SAI indexes on this same table, so
the fix stays inside Cassandra rather than adding the second system
you just avoided.
Neither of those statements is specific to Cassandra. They describe ANN retrieval wherever it runs, and teams building RAG on a dedicated vector store can reach the same two conclusions and apply the same two fixes. What Cassandra changes is the operational surface rather than the ranking: a column and an index on a table you already replicate, back up, and secure.
Native doesn’t mean equivalent. A dedicated search platform still wins on lexical ranking and relevance tuning, but it changes the starting assumption and depending on the use case, a second system doesn’t need to be a default to accept.
Apache Cassandra vector search version and operational notesVector search on Cassandra is continuing to mature, so treat patch releases as part of planning.
- Prefer 5.0.7 or later. Earlier
5.0 patches could return stale neighbors after overwrites and
tombstones, and
WHEREplus ANN queries were prone to timeouts. CASSANDRA-20086 reworked that path in 5.0.7 and 6.0. This project uses 5.0.9. Note that Instaclustr does not support 5.0.7 nor 5.0.8. - Cassandra 6.0 isn’t GA yet. Alpha1 arrived in March 2026, alpha2 in August. The work since 5.0 has been correctness rather than new capability, so nothing above changes.
- Cassandra 5.0.6 is Generally Available on the NetApp Instaclustr Managed Platform. This works for the demo, which is read-only against a freshly seeded table. Track the platform version list before putting vector writes into production.
Run the three questions against a live Cassandra 5 table on Docker. The demo application is from cassandra-vector-demo. Clone it, follow the README, and compare your rankings to the output in this series.
The same CQL runs on Instaclustr for Apache Cassandra 5.0. A new cluster takes a few clicks in the console; start a free trial and point the demo at the managed contact point instead of 127.0.0.1. You write the table; NetApp manages the cluster.
Further reading on native vectors: Apache Cassandra 5.0 vector search, similarity search with the vector type.
Part 2 of this series continues with the same Summit table and uses hybrid RAG to improve correctness.
The post Vector Search in Apache Cassandra for RAG Retrieval appeared first on Instaclustr.