How Sprig Replaced Postgres, ClickHouse & Redis…with 4-8x Better Latency
With ScyllaDB, a small engineering team could focus on building their product instead of battling their databases “Just use Postgres” is the logical path for a startup figuring out product-market fit. But Sprig – an AI-powered product research platform – hit that fit much sooner than expected. And the sheer scale of the resulting data (over 1.3T events and 75 B attributes from 15B visitors)… SourceHow Medium Powers Real-Time Recommendations at 1M OPS
Inside Medium’s move from relational features to list features in its ScyllaDB-based feature store “Keep readers reading” is the not-so-simple goal of Medium’s recommendations system. To predict what’s most likely to appeal to a particular reader at any given time, Medium continuously processes user activity signals (stories read, recommendations shown, follows, likes, etc.). SourceScyllaDB Customer Experience Spotlight: Susie Solis
Welcome to the latest installment of a new blog series introducing some of the experts you might encounter when you work with ScyllaDB. (Past ones: Tyler Denton, Faisal Saeed). Today we’re featuring Susie Solis, Technical Support Engineer on the Customer Experience team here at ScyllaDB. She lives in Colorado and has been at ScyllaDB for almost 3 years. Let’s learn a little about Susie… I’… SourceDDIA 2nd Edition Excerpt: On Scalability
Martin Kleppmann and Chris Riccomini’s scalability considerations for designing data-intensive applications The following is an excerpt from Chapter 2 of the shiny new Designing Data-Intensive Applications Second Edition, by Martin Kleppmann and Chris Riccomini. It is published here with permission from O’Reilly. For more background on the book (known as the bible of data systems) and its… SourceLessons Learned from Real-World NoSQL Database Migrations
Discover the strategies, challenges, and trade-offs teams faced in a few real-world migrations to ScyllaDB In “What Matters Most for NoSQL Migrations”, I shared my top strategies for planning, executing and de-risking a NoSQL database migration. I discussed key steps like schema and data migration, data validation and important considerations such as technology switches, tooling… SourceOffloading I/O to Dedicated Cores: An Asymmetric io_uring Backend for Seastar and ScyllaDB
We moved I/O off application cores to dedicated cores to give compute-heavy workloads more breathing room Modern high-throughput systems often run on machines with many cores. At that scale, one practical problem becomes increasingly visible: CPU time spent on low-level I/O handling competes directly with CPU time needed for application-level work. In this post, we explore how shifting more I/ SourceWhat Matters Most for NoSQL Migrations
How to prioritize the things that matter most for planning, executing and de-risking your NoSQL database migration No team wants to take on a database migration. But many teams are migrating NoSQL databases to optimize performance at scale, alleviate maintenance burdens and/or control costs. If a migration is on your team’s radar, how do you make the process as painless as possible… SourceBuild Durable Chat Memory for RAG Using ScyllaDB and LangChain
How to replace LangChain’s in-memory chat history with ScyllaDB — so your RAG chatbot retains context across restarts and scales across replicas This post demonstrates how to integrate ScyllaDB Vector Search into your LangChain project for RAG use cases, as well as how to use ScyllaDB as a durable conversation memory within LangChain. Large language models are trained on a fixed snapshot of… SourceScyllaDB Is Now Supported in MCP Toolbox for Databases
Connect your AI agents to ScyllaDB using the new ScyllaDB integration in MCP Toolbox for Databases ScyllaDB is now a supported database source in MCP Toolbox for Databases (as of v1.5.0). This means that your agents and applications can query your ScyllaDB clusters using the same framework they use to connect to Postgres, MySQL, BigQuery, and dozens of other databases. This post walks through… SourceAgent Memory at Monster Scale with Mem0 and ScyllaDB Cloud
Combine Mem0’s memory management with ScyllaDB’s persistence features to deploy large-scale AI agents Scaling AI agents to handle persistent memory for a large number of users (e.g. 500k+ DAU) introduces engineering bottlenecks. These complications generally compound across two areas: how context is filtered for the model and how that data is stored globally. Mem0 addresses context efficiency… SourceWhat’s Coming in Cassandra? Key Apache Cassandra CEPs to Watch
IntroductionApache Cassandra’s contributors continue to push the database forward, and the Cassandra Enhancement Proposal (CEP) process is where that work takes shape. A CEP is a proposal to design, discuss, and build a meaningful change, with the author signaling real intent to implement it and to gather community consensus along the way.
We have previously covered CEPs here, some of which are anticipated to be present in Apache Cassandra 6 (currently in alpha). In this article we look at five CEPs that together touch many layers of Cassandra: replica consistency (CEP-45: Mutation Tracking), data placement and balancing across a cluster (CEP-60: Flexible Placements), cluster administration (CEP-38: CQL Management API and CEP-62: Cassandra Configuration Management via Sidecar), and efficient use of the underlying hardware (CEP-49: Hardware-accelerated compression).
All of these CEPs have been accepted, but as with all open source development, inclusion in a future release depends on successful implementation, community consensus, testing, and approval by project committers.
CEPs Discussed- CEP-38: CQL Management API
- CEP-45: Mutation Tracking
- CEP-49: Hardware Accelerated Compression
- CEP-60: Flexible Placements
- CEP-62: Configuration Management via Sidecar
What it does: Adds a native Cassandra CQL interface so operators can run cluster administration tasks directly through CQL instead of depending on JMX-backed tooling.
Most Cassandra admin tasks, from taking a snapshot to running a compaction, run through JMX MBeans, and the tools operators depend on, like nodetool and the Cassandra Sidecar, all speak to them over JMX. The ecosystem has worked around this over the years by wrapping JMX in REST APIs or bypassing it with Java agents, but these layers sit on top of an internal API that was never designed as a stable contract. There are further drawbacks to that coupling, from JMX’s security exposure and operational complexity to the cost of maintaining nodetool and the lack of structured command metadata from the server.
CEP-38 intends to make CQL the management interface to run commands directly, removing the dependence on JMX and external tooling in addition to aligning administration with the same interface developers already use. Defining each command once in a single registry with structured metadata gives the agents and plugins that expose a REST API something solid to build on.
Here’s an example using the CQL syntax, per the CEP’s documentation:
EXECUTE COMMAND forcecompact WITH keyspace=distributed_test_keyspace AND table=tbl AND keys=["k4", "k2", "k7"];
Just as important, it moves command execution toward an asynchronous, observable model: instead of executing and blocking, a command can be submitted, return an identifier to track it, and have its result observed afterward.
This flow is intended to lay the groundwork for automation and higher-level workflows in the future. Underpinning both the CQL interface and this execution model is a single command registry that defines each command once and exposes it consistently across interfaces. This would prevent drift between JMX, the CLI, and any REST layer.
With behavior centralized, nodetool and cqlsh stop being separate implementations and become thin entry points over the same operations. A dedicated management surface that can be reached on its own admin port gives the control plane a clear boundary that higher-level orchestration can build on.
This CEP benefits operators administering clusters and developers building and working with management tooling. Crucially, the CEP doesn’t aim to remove or deprecate the existing MBeans or CLI tools. JMX will keep working, but it nudges the project toward a state where deprecating JMX could eventually become feasible.
CEP-45: Cassandra mutation tracking for replica consistencyWhat it does: Tracks individual writes by ID rather than comparing whole partitions across replicas.
Cassandra has two ways of catching writes that didn’t reach every replica, repair and read repair, but both work by pulling stored data from multiple nodes and comparing it, which is expensive.
Repair ships whole partitions between nodes when it finds a discrepancy, driving up streaming and compaction work and making very large partitions impractical. On the other hand, read repair only fixes the slice of a partition a query touched. This can leave a write half-applied and undermine Cassandra’s partition-level write atomicity. It also can’t provide read monotonicity—an important property of quorum reads and writes—for witness replicas without read-repairing nearly every read, which has limited their usefulness.
CEP-45 takes a different approach. Rather than comparing data on disk, each write is tracked individually: the coordinator stamps every write with a unique ID that travels to the replicas, and each replica records which IDs it has applied. At read time, one replica returns the data plus a summary of its applied IDs while the others return only that summary; matching summaries mean the data is accurate, and any gaps are filled by sending the specific missing writes. A background process continually reconciles these IDs across replicas and establishes a lower bound (kind of like a watermark) that signals older log entries can be cleaned up.
The feature is enabled per keyspace or table through a new replication-type setting and reuses Accord’s addressable commit log, which adds an index over Cassandra’s commit log so individual entries can be retrieved by ID.
Repair and read repair have long been operational burdens for Cassandra operators. Reconciling at the level of individual writes instead of whole partitions should cut streaming and compaction cost and ease partition-size limits.
CEP-49: Hardware-accelerated compressionWhat it does: Offloads compression work to hardware accelerators where available, freeing CPU for other tasks.
Cassandra ships with four compressors (LZ4, Zstd, Deflate, and Snappy) and compressing and decompressing data eats a meaningful share of CPU during flush and compaction. Compression can also apply to the commitlog and to data moving across the network.
Some newer processors carry built-in accelerators for this work, such as Intel’s QuickAssist Technology (QAT) on Intel Xeon chips, which can accelerate LZ4, Zstd, and Deflate. The proposal intends to hand compression off to that hardware where it exists, freeing CPU for other tasks and speeding up compression itself.
CEP-49 adds a framework that uses the accelerator when present and reverts to the software compressor otherwise, with room to plug in other accelerators later. Backends ship as separate plugins that Cassandra discovers at startup and it falls back to the standard compressor if a plugin fails.
The main beneficiaries are operators running compression-heavy workloads on capable hardware, though the framework is also designed to support other hardware-based compressors in the future.
Note: The hardware must already be configured correctly, and anything not functioning falls back to default software-based compression.
CEP-60: Flexible Placements for cluster scaling and balancingWhat it does: Decouples data placement from token ring position, enabling steadier cluster utilization and more granular scaling.
In the current model, a node’s token positions on the ring determine which data it owns, which causes several problems. Growing a cluster cheaply tends to require doubling it; while a node joins, its range is temporarily served by an extra replica, adding load. On vnodes, tokens can’t be moved, so fixing an unbalanced ring falls on the operator and there’s no way to plan a large change as one operation or break a long one into smaller retriable steps.
CEP-60 decouples placement from token position, making the “tablet”—a range for a specific keyspace/table pair—the unit of ownership, with replicas assigned directly. Built on CEP-21’s transactional cluster metadata, it lets bootstrap and streaming run in smaller resumable steps. It also decides where data lives using per-range load and capacity metrics, moving away from ring percentage.
A central benefit is better node density. Because the cluster stays close to balanced at any size, operators can run nodes at higher, steadier utilization. By contrast, token-based growth produces a sawtooth pattern, forcing operators to provision for the peak and pay for idle capacity. Keeping nodes near a target utilization, and growing a few nodes at a time, translates into fewer wasted machines and potentially lower cost.
CEP-62: Cassandra configuration management via SidecarWhat it does: Adds a Sidecar REST API for programmatically reading and modifying cassandra.yaml and JVM options files.
Many Cassandra settings, like memtable configuration, SSTable
options, and storage_compatibility_mode, live in
cassandra.yaml and can’t be changed while a node is
running. Additionally, startup tuning such as heap size and garbage
collection lives in JVM options files. Runtime settings can be
adjusted through JMX, but for these on-disk files Cassandra offers
no programmatic interface, leaving operators to edit them by hand
or with custom scripts. An earlier addition let the Cassandra
Sidecar start and stop instances, but it still couldn’t touch the
configuration those instances read at boot.
CEP-62 fills that gap with a Sidecar REST API for reading and
changing cassandra.yaml and JVM options. It layers a
sparse “overlay” of explicit changes on top of a base template and
merges the two into the configuration a node actually uses; a
pluggable provider can keep overlays locally or in a central system
like etcd or Consul. A version-aware check rejects settings a given
Cassandra version wouldn’t recognize, reducing the risk that a typo
or unsupported setting leaves a node unable to start. Changes apply
on the next restart, and everything lives in Sidecar with Cassandra
left untouched.
This helps operators managing configuration across many nodes, especially those wiring Cassandra into centralized configuration tooling. It’s disabled by default and purely additive, so existing deployments and anyone not running Sidecar are unaffected, and it lays groundwork for later work on driving Cassandra upgrades through Sidecar.
ConclusionAltogether, these proposals show a project investing in the things that matter most to the people who run it: reliability, operability, and efficiency. Mutation tracking and flexible placements aim to make data consistency and cluster scaling less costly and less manual. The CQL management API and Sidecar-based configuration management give operators stabler, more programmable ways to administer their clusters. Hardware-accelerated compression squeezes more out of modern hardware.
Each feature aims to lower the operational challenges of running Cassandra at scale both for self-hosted clusters and managed Cassandra providers such as NetApp Instaclustr, and several CEPs lay groundwork that future enhancements will build on.
Following the CEP process is one of the best ways to see where Cassandra is headed. We’ll keep tracking these proposals as they move through implementation, and we look forward to seeing them land in users’ hands in future releases.
Ready to run Cassandra without the operational complexity? Try NetApp Instaclustr for Apache Cassandra free for 30 days today! Our managed platform handles the infrastructure, configuration, and operational heavy lifting so your team can focus on building applications.
The post What’s Coming in Cassandra? Key Apache Cassandra CEPs to Watch appeared first on Instaclustr.
ScyllaDB vs Aerospike, Wide-Column vs. Key/Value
Wide-column flexibility doesn’t have to come at the expense of performance — see where the two models differ, where each one wins, and why you no longer have to choose Aerospike published a paid benchmark to show it’s faster. Color me surprised…it defined the winner as itself. Aerospike benchmarked the one workload its architecture is built for. This article shares the fuller picture: what a… SourceCutting P99 Latency 1000X During Connection Storms by Hardening ScyllaDB Admission Control
ScyllaDB successfully mitigated performance-degrading connection storms by optimizing caching, throttling, and password hashing to achieve a 1000x reduction in tail latency The story begins with a customer-visible problem: node restarts spiked P99 latency to 5 seconds for a full minute, while typical latency for the cluster was only 4 milliseconds. The cause was a flood of new connections. SourceHow ScyllaDB’s Trie-Based Index Delivers Up to 3X More Throughput
By transitioning from separate summary and index files to a prefix tree, we optimized cache efficiency, reduced disk I/O, and reduced memory overhead Trie-based SSTable index format was added in ScyllaDB 2025.4. Since then, it has evolved and matured to become the default index format in ScyllaDB 2026.2. In this post, we deep dive into the format change, present its pros and cons… SourceScyllaDB 2026.2: DynamoDB Streams and Vector Search, Trie Indexes, and Strongly Consistent Tables
ScyllaDB 2026.2 brings a combination of GA new features, exciting experimental features, and multiple stability and external use case improvements. The updates include: 2026.2 is more stable, works faster, and better than any past release. You are encouraged to upgrade to it, and use it for any new deployment. For the full release notes, see this forum post. Alternator is a native… SourceInstaclustr product update: June 2026
Here’s a roundup of the latest features and updates that we’ve recently released.
If you have any particular feature requests or enhancement ideas that you would like to see, please get in touch with us.
Major announcements AI Search for OpenSearch is now generally available on the NetApp Instaclustr Managed PlatformAI Search for OpenSearch is generally available on the NetApp Instaclustr Managed Platform. It brings semantic search, hybrid search, and retrieval-augmented generation (RAG) without the complexity of managing software, infrastructure, or operational management. General availability expands on the public preview, adding support for external LLM and embedding services such as Amazon Bedrock and OpenAI for enterprise search, e-commerce, support chatbots, and observability-style use cases. Unlock new possibilities with AI search—learn more.
Introducing Kafka Client Telemetry: Centralized client metrics for Instaclustr Managed Apache Kafka®NetApp is introducing Client Telemetry for Instaclustr for Apache Kafka®, designed to deliver broker-integrated visibility into Kafka client and application-level metrics, with telemetry export and centralized collection. Instaclustr for Apache Kafka users can gain visibility into client behavior such as connection status, request rates, error rates, and latency from the broker, simplifying monitoring and supporting a holistic view of client interactions. Compliant Kafka clients collect metrics and push them to the brokers; brokers use an OpenTelemetry Collector to forward metrics to a customer-specified destination, with Prometheus 3.0+ and Datadog supported in this initial release.
Powering low-latency analytics with ClickHouse® and Amazon FSxInstaclustr Managed ClickHouse integrated with Amazon FSx for NetApp ONTAP is built to run analytical queries directly on file-based data that can transparently tier to lower-cost capacity, without relying on extra staging layers, ingestion pipelines, or format-specific copies to make data queryable. The integration now supports deployments where compute and storage can reside in different VPCs or AWS accounts, enabling flexible, enterprise-grade architectures with consistent storage access across network and account boundaries.
Other significant changes Apache Cassandra®- Self-service iccassandra password reset — customers can now reset their iccassandra database password directly from the console via the Connection Info page, eliminating the need to raise a support ticket. The new password is displayed for 5 days before being automatically removed.
- Released Apache Cassandra v4.1.10 into General Availability on the NetApp Instaclustr Managed Platform, delivering a stability-focused patch release, while deprecating Apache Cassandra 4.1.9.
- Kafka and Kafka Connect 3.9.2 released to General Availability.
- Kafka and Kafka Connect 4.1.2 released to General Availability.
- Karapace Schema Registry 5.2.0 and Karapace Rest Proxy 5.2.0 are added support for Kafka clusters.
- ClickHouse v25.8.24 released to General Availability.
- New c7g.8xlarge node size on the AWS provider has been added to support OpenSearch clusters.
- OpenSearch 3.5.0 released to General Availability.
- AI Search is now available on the free trial.
- PostgreSQL 18.3, 17.9, and 16.13 and PgBouncer 1.25.1 released to General Availability.
- The new AWS region, ap-southeast-6 (New Zealand), has been added.
- Cluster tag management improvements — multiple enhancements to tag search, display, and validation in the console and API, including prevention of duplicate tag keys for better data consistency.
- We’re preparing to introduce GPU nodes for OpenSearch on the NetApp Instaclustr Managed Platform, bringing dedicated machine learning capabilities directly into your managed clusters. With GPU nodes, vector indexing can be up to 10x faster and CPU load is reduced, freeing cluster capacity for mission-critical workloads. Additionally, GPUs offer superior cost-efficiency compared to traditional CPU-based vector indexing, driving down the total cost of ownership.
- We’re close to launching PostgreSQL® integrated with FSx for NetApp ONTAP (FSxN) into GA, now including NVMe support—designed to deliver improved throughput, up to 20% observed greater throughput than we achieved with our public preview. This enhancement combines enterprise-grade PostgreSQL with FSxN’s scalable, cost-efficient storage for better cost, performance, and flexibility, while enabling ONTAP snapshots for backups, mirroring, and multi-region recovery—fast snapshot/restore and daily backups for large databases.
- NetApp Instaclustr plans to release the Remote MCP Gateway Service powered by AgentGateway on the Instaclustr Managed Platform. This service will let you, in minutes, provision and configure a production-ready Model Context Protocol gateway to provide LLM access to databases, application data infrastructure services, and REST APIs.
- Coming soon, NetApp Instaclustr will be launching the
Self-Service Bring Your Own Cloud (BYOC) feature for AWS, offering
a fully guided onboarding experience that allows customers to
connect their AWS accounts and begin deploying managed clusters
directly from the console — making it faster and easier for
customers who prefer to run clusters in their own cloud
environments.
Cluster DNS will soon be available for Apache Cassandra and Apache Kafka clusters on AWS allowing you to connect to your applications using simple, stable hostnames instead of long lists of IP addresses. When node IPs change due to scaling, replacement, or maintenance there is no longer a need to update client configuration.
- Need an end-to-end pattern for streaming analytics on AWS? The same-day three-part series How to build a streaming analytics pipeline with Terraform and Instaclustr, Part 1: Setting up your first Kafka® cluster, Part 2: Designing the complete data pipeline, and Part 3: Integrating with AWS VPC show how to stand up Kafka with Terraform, connect ClickHouse and Kafka Connect into a real pipeline, and finish with VPC integration for secure networking. Together the posts bridge provisioning, data flow design, and cloud networking without skipping the glue work that usually stalls proof-of-concepts.
- Apache Kafka 4.1.0 introduces the Streams Rebalance Protocol in early access for Kafka Streams: a broker-driven assignment model that eliminates client-side coordination, reduces “stop-the-world” rebalance pauses, and delivers smoother task assignment as Streams applications scale horizontally. For a walkthrough of when you need it, how to enable it, and what to expect, see What’s new in Kafka® 4.1.0? Introducing the new Streams Rebalance Protocol.
- OpenSearch 3.6 release bundles a wide set of upstream changes: ML Commons AI agent improvement such as token usage tracking, k-NN vector search performance improvements including Lucene Better Binary Quantization, Dashboards updates across AI chat and Explore, and OpenSearch APM for observability. For a single walkthrough of those themes, see OpenSearch version 3.6 release: smart agents and fast search. We’re currently testing OpenSearch 3.6 for compatibility and security purposes. Keep an eye on our release blog for more information about when this exciting new release will be available on the managed platform.
If you have any questions or need further assistance with these enhancements to the Instaclustr Managed Platform, please contact us.
SAFE HARBOR STATEMENT: Any unreleased services or features referenced in this blog are not currently available and may not be made generally available on time or at all, as may be determined in NetApp’s sole discretion. Any such referenced services or features do not represent promises to deliver, commitments, or obligations of NetApp and may not be incorporated into any contract. Customers should make their purchase decisions based upon services and features that are currently generally available.
The post Instaclustr product update: June 2026 appeared first on Instaclustr.