Building a Faster (Rust-Based) Python Driver for ScyllaDB

Binding Rust to Python with PyO3, zero-copy deserialization with yoke Since Python is one of the most commonly used languages for interacting with a database, having a robust ScyllaDB Python driver is essential. The first ScyllaDB Python driver was created ages ago, back when ScyllaDB was in the 2.x stage, and it’s been used ever since. As part of the ZPP Project at the University of Warsaw… Source

MCP Gateway RBAC tutorial: Add Role-Based Access Control to your AI agent with Auth0 and Apache Cassandra

This is a follow-up to “MCP Gateway tutorial: Connect an AI Agent to Apache Kafka and HTTP Server Backends,” extends the original support chatbot with Auth0-based login, role-based access control (RBAC), and an Apache Cassandra backend.

In this MCP Gateway RBAC tutorial, you will configure MCP Gateway to enforce role-based access control for an AI agent. The walkthrough adds Auth0 authentication, maps user roles to MCP Gateway Access Control Lists, connects an Apache Cassandra backend, and shows how each persona receives only the tools their role is allowed to use.

The goal is to support least-privilege access for AI agent tools at the gateway layer: Auth0 authenticates the user, MCP Gateway maps the user’s role to an Access Control List, and the agent receives only the tools that role is permitted to call.

Get the code: MCP-Gateway/shoe-store-support-chatbot-rbac

Recap: where we left off

The previous article built a support chatbot for a shoe e-commerce company, backed by one MCP (Model Context Protocol) Gateway Virtual Server (supportchatbot) with two backends: ordersapi (HTTP) providing get_orders, and supporttickets (Apache Kafka) providing submit_request. That version had a single, shared set of tools for every user.

This article adds Auth0 login and three distinct roles (support-agent, merchandiser, and readonly-auditor), plus a third backend type, Cassandra, so each persona gets exactly the access their job calls for, configured entirely at the MCP Gateway.

Prerequisites

From the previous tutorial, you should already have:

  • An MCP Gateway Virtual Server (supportchatbot).
  • The ordersapi (HTTP) backend with a get_orders tool.
  • The supporttickets (Kafka) backend with a submit_request tool.

Everything involving Auth0, OAuth, roles, and Access Control Lists (ACLs) is new in this article.

Tooling: Python 3.11+, the uv package manager, and an AWS account with Amazon Bedrock model access enabled for the model you’re using, in the region you’re calling (e.g. us-east-1).

The use case: a product catalog with tiered cost visibility
  • support-agent (customer-facing staff) gets product info (name, price, stock) to help shoppers, plus order lookups and ticket filing.
  • merchandiser (buying/purchasing staff) gets supplier cost and margin data to negotiate pricing, commercially sensitive information relevant to their role, distinct from customer-service tooling.
  • readonly-auditor gets full read visibility across both domains for review purposes, with action (filing a ticket) reserved for support-agent.
What is Role-Based Access Control (RBAC), and why does it matter for an AI agent?

Role-Based Access Control means a user’s identity maps to a fixed role, and every access decision is checked against an explicit allow-list. In MCP Gateway, that role determines which tools an AI agent can discover and call, so each persona gets exactly the access their job requires: no more and no less.

This matters especially for an LLM-backed agent, for three reasons:

  • It keeps every session’s capabilities predictable, no matter how the conversation goes. The chatbot handles free-text customer input, and however that conversation unfolds, a support-agent session simply can’t reach catalogdata_get_supplier_cost: it’s not part of that role’s tool set, so there’s nothing to accidentally or deliberately reach for. The boundary is a property of the session, not a judgment call the model has to get right in the moment.
  • It lets you express two independent things: what a role can see, and what it can do. merchandiser can see cost data but can’t file tickets; readonly-auditor can see everything but can’t act on any of it. Splitting visibility from action this way maps naturally onto how organizations already assign responsibility.
  • It moves enforcement to a layer that’s easy to reason about. Rather than relying on prompt instructions like “don’t answer cost questions for this role,” which work well in the common case but depend on the model’s judgment every time, the MCP Gateway checks the ACL before a tool call is honored, giving you one clear, auditable place where access is decided.
Architecture: Connecting Auth0, MCP Gateway RBAC, and Cassandra tool access

MCP gateway blog supporting image 1

Figure 1. Routing role-based AI agent requests to HTTP, Kafka, and Cassandra backends through the MCP Gateway

Step 1: Provision the Cassandra cluster

In the Instaclustr console, Create Cluster → Apache Cassandra, choosing a data center/region and node size appropriate for a demo.

Once the cluster reaches Running, note its Data Center ID from the Details page for Step 3.

Use the console’s Connection Info tab and default superuser credentials to connect via cqlsh for the next step.

Step 2: Set up the Cassandra schema, seed data, and a least-privilege role
CREATE KEYSPACE IF NOT EXISTS catalog WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 3}; CREATE TABLE catalog.products ( productid text PRIMARY KEY, name text, description text, category text, listprice decimal, stockstatus text ); CREATE TABLE catalog.supplier_costs ( productid text PRIMARY KEY, supplierid text, unitcost decimal, marginpercent decimal, suppliercontact text ); CREATE ROLE mcpgateway_catalog WITH PASSWORD = '<password>' AND LOGIN = true; GRANT SELECT ON catalog.products TO mcpgateway_catalog; GRANT SELECT ON catalog.supplier_costs TO mcpgateway_catalog; INSERT INTO catalog.products (productid, name, description, category, listprice, stockstatus) VALUES ('SHOE-001', 'Midnight Runner', 'Black low-top running shoe', 'running', 129.99, 'in_stock'); INSERT INTO catalog.products (productid, name, description, category, listprice, stockstatus) VALUES ('SHOE-002', 'Trailblazer Boot', 'Brown waterproof hiking boot', 'hiking', 189.99, 'backordered'); INSERT INTO catalog.supplier_costs (productid, supplierid, unitcost, marginpercent, suppliercontact) VALUES ('SHOE-001', 'SUP-88', 42.00, 67.7, 'orders@northstarfootwear.com'); INSERT INTO catalog.supplier_costs (productid, supplierid, unitcost, marginpercent, suppliercontact) VALUES ('SHOE-002', 'SUP-14', 71.50, 62.4, 'sales@trailgearco.com');

Connect with the credentials and endpoint from your cluster’s Connection Info tab before running this (see Step 1). mcpgateway_catalog is read-only and scoped to just these two tables: a second, defense-in-depth layer of access control underneath the MCP Gateway’s own ACLs.

Notes:

  • Identifiers are written lowercase throughout, matching how Cassandra represents unquoted identifiers internally: table/column names and named bind markers alike. Writing them lowercase from the start keeps the CQL here consistent with the tool configuration in Step 3.
  • Remember to replace <password> in the CREATE ROLE mcpgateway_catalog statement with a real, strong password before running this; it’s a placeholder, not a literal value. Keep it somewhere secure, since you’ll need to re-enter it in Step 3 when you create the catalogdata backend on the MCP Gateway (that’s where these credentials actually get used).
Step 3: Create the Cassandra backend and tools on the MCP Gateway

On the supportchatbot MCP Virtual Server → Create Backend → name catalogdata, type Cassandra, using the Data Center ID from Step 1 and the mcpgateway_catalog credentials from Step 2.

MCP gateway blog supporting image 2

Figure 2. The catalogdata Cassandra backend, showing its Cassandra Data Center ID, Cassandra Username (mcpgateway_catalog), and its two configured tools.

Open the new catalogdata backend and, under Tools, click Create Tool to create the below two tools.

get_product_details:

SELECT name, description, category, listprice, stockstatus FROM catalog.products WHERE productid = :productid
  • Parameter: productid (string)
  • Result fields: name (string), description (string), category (string), listprice (string), stockstatus (string)
  • Description: “Look up a product’s customer-facing details by product ID, including its name, description, category, list price, and current stock status.”

get_supplier_cost:

SELECT supplierid, unitcost, marginpercent, suppliercontact FROM catalog.supplier_costs WHERE productid = :productid
  • Parameter: productid (string)
  • Result fields: supplierid (string), unitcost (string), marginpercent (string), suppliercontact (string)
  • Description: “Look up internal supplier cost and margin data for a product by product ID. This is sensitive commercial data; only use it when explicitly asked, and only if you have access to this tool.”

MCP gateway blog supporting image 3

Figure 3. The get_product_details tool definition, showing its CQL query, productid parameter, and result fields.

Notes:

  • Declare Result Field Type as string for any Cassandra decimal column.
  • Give every tool a real, descriptive description. Beyond helping the model choose the right tool, AWS Bedrock’s Converse API expects every tool in the list to have one, so it’s worth treating as a required field alongside the CQL and parameters.
  • Apache Cassandra forces unquoted bind variable names to lowercase: :productId is interpreted internally as :productid. That’s why every identifier in this tutorial (table/column names and bind markers alike) is written lowercase from the start, keeping the CQL here consistent with the tool configuration above. If you ever need a case-sensitive parameter name, quote it in the query instead, e.g. :"productId".
Step 4: Configure Auth0: application, roles claim, and personas

4a. Application and Post-Login Action: Follow the Auth0 example configuration guide to:

  • Create a Regular Web Application and configure its Allowed Callback URIs / Web Origins.
  • Create the Post-Login Action that adds a roles claim when the mcp_roles scope is requested, add the MCP_GATEWAY_CLIENT_ID secret, deploy it, and apply it on the post-login trigger.
  • Note the audience (API Identifier), issuer, and JWKS URI as described there; needed in Step 5.

4b. Create the three Auth0 Roles: User Management → Roles → Create Role, once each for:

  • support-agent
  • merchandiser
  • readonly-auditor

4c. Create three test users, one per role: User Management → Users → Create User (or reuse existing test users), then on each user’s Roles tab, Assign Roles with the matching role from 4b.

Note: Keep role names consistent, case-for-case, across three places: the Auth0 Role, the user_roles claim value it produces, and the MCP Gateway ACL’s Role name field in Step 6. Matching these exactly is what makes each persona’s tool list show up correctly.

Step 5: Enable OAuth on the MCP Virtual Server

On the supportchatbot MCP Virtual Server’s OAuth configuration section (refer to the documentation for more details), enter the values noted in Step 4a:

  • Issuer: https://<domain>/
  • JWKS URI: https://<domain>/.well-known/jwks.json
  • Audience: your API Identifier
  • Scopes Supported: mcp_roles
  • Roles Claim Name: https://mcp_gateway/user_roles

Save/apply.

MCP gateway blog supporting image 4

Figure 4. The supportchatbot MCP Virtual Server’s OAuth configuration with all three backends.

Step 6: Create the three Access Control Lists

On the MCP Gateway’s ACL screen, you pick each tool by its bare name; the Backend column disambiguates which tool you mean when the same name could theoretically exist on more than one backend:

support-agent

Tool Backend get_orders ordersapi submit_request supporttickets get_product_details catalogdata

readonly-auditor

Tool Backend get_orders ordersapi get_product_details catalogdata get_supplier_cost catalogdata

merchandiser

Tool Backend get_product_details catalogdata get_supplier_cost catalogdata

MCP gateway blog supporting image 5

Figure 5. The support-agent Access Control List, showing its explicitly allowed tools.

Each role has a distinct tool set tailored to its purpose: merchandiser focuses entirely on catalog data, while readonly-auditor sees everything the other two combined can see, with action reserved for support-agent.

Once exposed through the supportchatbot MCP Virtual Server, the agent sees each tool under a backend-prefixed name (e.g. catalogdata_get_product_details) to disambiguate tools coming from different backends; that’s the name that shows up in the app’s “Tools discovered” sidebar later, and in Bedrock’s tool-calling requests, even though you select it by its bare name here.

MCP gateway blog supporting image 6

Figure 6. The completed supportchatbot MCP Virtual Server, showing its OAuth configuration, all three backends, and the three role-based Access Control Lists.

Why the chatbot’s code needs no changes?

None of this required touching the chatbot’s code, because it was built to discover tools dynamically rather than hardcode them: mcp_manager.py‘s connect() attaches the user’s access token as a Bearer header once, at connection time, and every session it opens rides on that same authenticated connection for the rest of the login; get_tool_specs() then simply asks the MCP Gateway which tools this identity can see and turns whatever comes back into Bedrock’s toolConfig.tools format, with the MCP Gateway itself deciding (via the caller’s role and its Access Control List) which tools are included; and app.py‘s sidebar just renders that same list, grouping tool names by server with nothing hardcoded. So, when catalogdata‘s two tools were added at the MCP Gateway, they simply started showing up for the roles allowed to see them, with zero code changes anywhere in this repo.

Running the application

The full code for this tutorial is available at MCP-Gateway/shoe-store-support-chatbot-rbac in the instaclustr/code-samples repo.

cp config.example.json config.json # fill in your Auth0 tenant + MCP Gateway URL uv sync export AWS_ACCESS_KEY_ID="..." export AWS_SECRET_ACCESS_KEY="..." export AWS_DEFAULT_REGION="us-east-1" uv run streamlit run src/rbac_shoe_bot/app.py

This opens the app at http://localhost:8501. Compared to the shared-access version from the previous tutorial, config.json now includes a populated auth0 section, and the app opens on a Log in with Auth0 screen.

Click Log in with Auth0 and sign in as one of your test users. To try a different persona, click Log out and log in again; each login prompts for credentials fresh, so you can move between test users freely.

MCP gateway blog supporting image 7

Figure 7. A readonly-auditor session in the chatbot, showing its discovered tools in the sidebar and a successful get_supplier_cost lookup for SHOE-002.

Notes:

  • Enable model access for bedrock.model_id in the Bedrock console for your region ahead of time; it’s a quick, one-time opt-in per model/region, and worth doing before your first run.
  • The sidebar’s Tools discovered list is the fastest way to confirm your ACL setup is working as intended; a merchandiser session should show only the two catalogdata_* tools.

Demo script

Using the two seeded products:

  • support-agent: “What’s the price and stock status of product SHOE-001?” → succeeds. “What do we pay our supplier for it?” → no tool available for this role.
  • merchandiser: “What’s our cost and margin on SHOE-001?” → succeeds. “Show me my recent orders.” → no tool available for this role.
  • readonly-auditor: sees both cost and orders, with ticket filing reserved for support-agent.

Note: These prompts reference product IDs directly rather than descriptions like “black shoes,” since get_product_details looks up by exact productid. Referencing known IDs keeps the demo focused on RBAC rather than search.

Why this matters beyond the demo

Authentication, three distinct roles, and a whole new backend technology were all added here through configuration: Auth0 setup, MCP Gateway OAuth config, ACLs, and the Cassandra backend/tools, without changing the chatbot’s own code. That’s the strength of centralizing access control at the MCP Gateway layer: the same access rules apply consistently, however a conversation unfolds, because the boundary lives outside the agent’s own reasoning.

Try it yourself

If you want to add role-based access control to your own AI agent, you can spin up MCP Gateway and an Apache Cassandra cluster today: start a free Instaclustr trial, no credit card required. From there:

If you’d rather start from the simpler, single-role version, the previous tutorial walks through connecting an agent to Kafka and HTTP backends without the auth layer.

The post MCP Gateway RBAC tutorial: Add Role-Based Access Control to your AI agent with Auth0 and Apache Cassandra appeared first on Instaclustr.

Vector Search in Apache Cassandra for RAG Retrieval

Semantic retrieval using one Cassandra table to understand how it works, where it shines, and what to watch for.

Vector search in Apache Cassandra ® lets teams store embeddings, index them with Storage-Attached Indexing, and run approximate nearest neighbor queries directly in CQL without managing a separate vector database. For RAG retrieval, that means Cassandra can act as both the operational database and the semantic search layer when the retrieval problem is primarily meaning-based.

Search used to require the person asking to already know the answer’s vocabulary. A technician precisely asks, “Recall 24V-113: tailgate latch”; a truck owner asks, “Is there a recall on the tailgate?” Embeddings close that gap by mapping meaning into geometry, so similar ideas land near each other in vector space regardless of the words used. That geometry is the retrieval layer underneath most AI applications. Using retrieval-augmented generation (RAG), a pattern for answering questions from supplied data instead of the model training set alone, the chatbot can be instructed to only answer from the documents vector search hands it. As a result, the retrieval quality sets the ceiling on the answer.

The standard approach utilizes a dedicated vector store alongside the operational database. While Cassandra is not the first operational database to keep embeddings next to the data they describe, version 5.0 made it a native part of the data model. When the use case fits, querying embeddings on the same infrastructure that houses the data can cut both complexity and cost: one database to operate instead of two, and no application code responsible for keeping them in sync.

This is Part 1 of 2. Part 1 demonstrates semantic retrieval against a realistic knowledge base containing vehicle recalls, tow ratings, and warranty terms using Cassandra. Then, we show where Approximate Nearest Neighbor (ANN) alone is not the complete answer, which sets up a different approach, hybrid RAG, in Part 2.

What is vector search on Apache Cassandra?

Vector search on Apache Cassandra lets you store, index, and query high-dimensional embeddings natively in the database using the vector<float, N> column type and Storage-Attached Indexing. Introduced in Cassandra 5.0, it enables ANN search through an ORDER BY … ANN OF clause in CQL, with no separate vector store.

At write time, title + body is embedded and stored in a vector column on the same row as the text. At query time the question is embedded with that same model, and Cassandra is asked for the nearest stored vectors. Three pieces of CQL do the whole job: declare the column, attach an SAI index so ANN is possible, and query with ORDER BY … ANN OF plus an optional similarity score in the select list.

CREATE TABLE docs ( id text PRIMARY KEY, title text, body text, category text, model text, model_year int, embedding vector<float, 384> ); CREATE INDEX ON docs (embedding) USING 'sai' WITH OPTIONS = {'similarity_function': 'COSINE'}; SELECT id, title, category, model, model_year, similarity_cosine(embedding, ?) AS score FROM docs ORDER BY embedding ANN OF ? LIMIT 5;

Each piece of that, in turn:

  • vector<float, N> is a native CQL column type holding an embedding of N dimensions, sitting on the row it describes rather than in another system. It’s 384 here because all-MiniLM-L6-v2 (which is used for the demo application) produces 384-dimensional vectors; document and query vectors must always use the same model and dimension.
  • SAI (Storage-Attached Indexing) attaches a searchable index to the table as rows are written. For vector columns it plugs in JVector, an ANN engine that walks a graph of nearby embeddings for fast top-k retrieval.
  • similarity_function accepts COSINE (the default), DOT_PRODUCT, or EUCLIDEAN. Changing it later means dropping and recreating the index.
  • ORDER BY embedding ANN OF ? returns approximate nearest neighbors ordered by similarity. This is fast, but not a guarantee of the true global nearest neighbor.
  • similarity_cosine(embedding, ?) is optional and exists to make the ranking legible. It returns a 0-to-1 value rather than raw cosine, so 0.5 means a document unrelated and a score of 0.75 is not “three-quarters relevant.”

Writes don’t need special handling because the embedding binds as one more parameter on an ordinary prepared INSERT alongside the document, and SAI indexes it as the row is written, under the replication factor and consistency level you already use. In addition, vectors impose no special replication requirements. The demo application’s keyspace is SimpleStrategy with RF 1 only because it’s a single-node Docker cluster.

Setting the scene

Throughout this series, you own a fictional 2024 Summit 1500 and you’re asking a support bot a question. The knowledge base is fictional but realistic: 21 documents covering 12 recalls, tow ratings, payload figures, maintenance intervals, and warranty terms for Summit 1500 and 2500 pickups.

The recall documents matter more than their count suggests. They share one sentence skeleton, mimicking how a real recall notice would read, making them a clean test of something specific: whether an embedding can resolve an identifier when the surrounding prose is nearly identical.

Everything in this series runs on Apache Cassandra 5.0.9 in Docker, with embeddings generated locally from the all-MiniLM-L6-v2, an embedding model, so there’s no API key to supply.

Companion app: cassandra-vector-demo. Follow the instructions in the README to reproduce the output below. Running the application isn’t required to understand the concepts in this series.

How Apache Cassandra vector search fits into a RAG pipeline

Throughout this demo, you never directly query Cassandra. Rather, you’ve asked a question, expect an answer back, and never see the retrieved document’s similarity scores. Vector search is the middle step that chooses which rows the language model is allowed to read:

  1. The application takes the question as text.
  2. The same embedding model used at insert time turns that question into a vector.
  3. Cassandra returns a small list of nearest documents (ORDER BY embedding ANN OF … LIMIT k).
  4. The application copies those titles and bodies into a prompt, next to the original question. That block is the context.
  5. The language model writes an answer using (in a well-behaved RAG setup) only that context. It doesn’t scan the knowledge base, only the handful of chunks retrieval picked.
Apache Cassandra vector search examples: Paraphrase, exact IDs, and vague queries

We run three Summit owner questions against the table above and read what Cassandra ranks. The demo doesn’t call an LLM after retrieving the results; the point is whether they are trustworthy before an LLM ever sees it. The first question is a paraphrase, which is what embeddings exist to solve. The other two are the limits you plan around: an exact identifier, and a question too vague to have one answer. Both are properties of similarity search rather than of Cassandra; the same model against Pinecone, pgvector, or OpenSearch kNN should produce the same or similar ordering and both are fixed with a query change rather than a database change. The output below is generated from src/demo.py against Cassandra 5.0.9 with all 21 documents seeded.

Outcome 1: Everyday paraphrase finds the right recall

“Is there a recall on the tailgate?”

python src/demo.py "Is there a recall on the tailgate?"QUERY: Is there a recall on the tailgate? 1. [recalls] recall-24v-113 [1500 2024] Recall 24V-113 tailgate latch score=0.7767 2. [recalls] recall-23v-011 [1500 2023] Recall 23V-011 instrument cluster score=0.7032 3. [recalls] recall-23v-088 [1500 2023] Recall 23V-088 backup camera score=0.6935 4. [recalls] recall-23v-145 [2500 2023] Recall 23V-145 brake booster score=0.6934 5. [recalls] recall-22v-078 [2500 2022] Recall 22V-078 trailer brake module score=0.6808

Vector search ranks Recall 24V-113 (tailgate latch) first, and the margin is wide: 0.7767 against 0.7032 for the runner-up. Nobody wrote the word “tailgate” into the question expecting a latch recall, and the model bridged it anyway, against 11 other recall documents written in nearly identical prose. Everyday paraphrase is exactly what embeddings are for, and one CQL query against one table is the whole implementation.

Outcome 2: An exact recall ID lands nowhere near the top

“What is recall 24V-330?”

Now you have the recall number.

python src/demo.py "What is recall 24V-330?" QUERY: What is recall 24V-330? 1. [recalls] recall-23v-011 [1500 2023] Recall 23V-011 instrument cluster score=0.7871 2. [recalls] recall-24v-207 [1500 2024] Recall 24V-207 fuel pump relay score=0.7701 3. [recalls] recall-23v-088 [1500 2023] Recall 23V-088 backup camera score=0.7623 4. [recalls] recall-24v-113 [1500 2024] Recall 24V-113 tailgate latch score=0.7560 5. [recalls] recall-24v-155 [2500 2024] Recall 24V-155 glow plug controller score=0.7514

The document for 24V-330 sits at rank 9, scoring 0.7332, below eight notices that never mention it. The recall notices share one sentence skeleton, so the ID is a handful of characters inside an otherwise near-identical embedding, and ANN ranks by overall closeness. Cosine distance has no notion of an exact token match, which is why the same model produces this ranking in any vector store, and why the fix is a lexical path beside ANN rather than a better embedding model. Part 2 of the series works to address this limitation.

Outcome 3: A question retrieval cannot answer as asked

“How much can the Summit 1500 tow?”

python src/demo.py "How much can the Summit 1500 tow?" QUERY: How much can the Summit 1500 tow? 1. [towing] summit-1500-2024-tow [1500 2024] Tow ratings by configuration, 2024 model year score=0.8990 2. [towing] summit-1500-2023-tow [1500 2023] Tow ratings by configuration, 2023 model year score=0.8853 3. [towing] summit-2500-2024-tow [2500 2024] 2024 Summit 2500 heavy duty towing capacity score=0.8791 4. [specs] summit-1500-2024-payload [1500 2024] 2024 Summit 1500 payload capacity score=0.8276 5. [towing] hitch-classes Hitch receiver classes and weight distribution score=0.7550

The ranking looks reasonable until you read the bodies. The top four sit within 0.08 of each other and answer four different questions: 11,300 lb for the 2024 1500, 9,100 lb for the 2023 1500, 17,600 lb for the 2500, and 1,750 lb of payload, which isn’t a tow rating at all.

Retrieval did its job since all four are genuinely relevant to what was asked. The question just never supplied a model year, so nothing in the index can pick a winner. The fix is a WHERE model_year = ? alongside ANN, which is exactly why those columns are on the table, and will be utilized in Part 2.

When to use Cassandra vector search in production RAG applications

Vector search is enough when the question is about meaning and an incorrect response is recoverable: everyday paraphrase like “is there a recall on the tailgate,” similar tickets, or similar products already sitting in Cassandra. It isn’t enough when an identifier, a model year, or a customer record has to be exact, and no similarity threshold will separate those two cases for you. Those need a lexical path or a WHERE clause next to ANN, which is the work of Part 2 using additional SAI indexes on this same table, so the fix stays inside Cassandra rather than adding the second system you just avoided.

Neither of those statements is specific to Cassandra. They describe ANN retrieval wherever it runs, and teams building RAG on a dedicated vector store can reach the same two conclusions and apply the same two fixes. What Cassandra changes is the operational surface rather than the ranking: a column and an index on a table you already replicate, back up, and secure.

Native doesn’t mean equivalent. A dedicated search platform still wins on lexical ranking and relevance tuning, but it changes the starting assumption and depending on the use case, a second system doesn’t need to be a default to accept.

Apache Cassandra vector search version and operational notes

Vector search on Cassandra is continuing to mature, so treat patch releases as part of planning.

  • Prefer 5.0.7 or later. Earlier 5.0 patches could return stale neighbors after overwrites and tombstones, and WHERE plus ANN queries were prone to timeouts. CASSANDRA-20086 reworked that path in 5.0.7 and 6.0. This project uses 5.0.9. Note that Instaclustr does not support 5.0.7 nor 5.0.8.
  • Cassandra 6.0 isn’t GA yet. Alpha1 arrived in March 2026, alpha2 in August. The work since 5.0 has been correctness rather than new capability, so nothing above changes.
  • Cassandra 5.0.6 is Generally Available on the NetApp Instaclustr Managed Platform. This works for the demo, which is read-only against a freshly seeded table. Track the platform version list before putting vector writes into production.
Try it yourself

Run the three questions against a live Cassandra 5 table on Docker. The demo application is from cassandra-vector-demo. Clone it, follow the README, and compare your rankings to the output in this series.

The same CQL runs on Instaclustr for Apache Cassandra 5.0. A new cluster takes a few clicks in the console; start a free trial and point the demo at the managed contact point instead of 127.0.0.1. You write the table; NetApp manages the cluster.

Further reading on native vectors: Apache Cassandra 5.0 vector search, similarity search with the vector type.

Part 2 of this series continues with the same Summit table and uses hybrid RAG to improve correctness.

The post Vector Search in Apache Cassandra for RAG Retrieval appeared first on Instaclustr.

Building a Faster C# Driver for ScyllaDB…on our Self-Baked C#-Rust FFI Framework

How we accidentally wrote our own FFI framework when developing a driver – a wild ride into modernizing a decade-old legacy codebase while dodging undefined behaviour and chasing maximum throughput. In our previous post, we already explained how we forged our own FFI framework in the fires of inter-language alchemy to connect the managed world of C# with the highly performant Rust ecosystem. Source

Bringing QUIC to Seastar

TL;DR For a University of Warsaw student project in collaboration with ScyllaDB, we built a working QUIC transport for Seastar (the engine that powers ScyllaDB). We implemented it on top of the sans-I/O ngtcp2 library and adapted Seastar’s RPC layer to run over it. The benchmarks show a bounded, predictable cost on a lossless loopback, and a clear advantage once the network starts dropping packets. Source

Full-Text Search, Object Storage Backend, and More in ScyllaDB 2026.3

The latest updates should help you move even more workloads to ScyllaDB, at a fraction of the cost ScyllaDB’s latest release adds major updates like full-text search, object store backend (preview), Oracle Cloud Infrastructure integration, and multiple bug fixes. These new updates should help you move even more workloads to ScyllaDB, at a fraction of the cost. The updates include: 2026.3… Source

A Self-Baked Async FFI Framework for Rust C# Interop

How we got tokio and .NET’s async runtime talking to each other, over the C ABI There’s just one ScyllaDB, but there are plenty of ScyllaDB Drivers… For every language you want to write your application in, you need a driver to talk to ScyllaDB. Maintaining and developing the whole herd of drivers has been a tough task for us, the Driver Team. One day we came up with an intriguing idea… Source

Building a New Rust Driver for ScyllaDB’s DynamoDB API – with 58% More Throughput

How our new Rust driver load-balances DynamoDB-style requests across a ScyllaDB cluster, and how we extended Latte to measure its performance For years, Rust developers using ScyllaDB’s DynamoDB-compatible API (Alternator) had one option: use the standard AWS DynamoDB SDK. This interface was designed for a managed service behind a single endpoint – so every request lands on one node while the… Source