All articles

Managed Cassandra vs Self-Hosted in 2026

I’ve been using Cassandra since 2008 and can understand why a team would want someone else to handle its maintenance. There is plenty to do between keeping repairs on schedule and testing recovery without also having to maintain scripts for every routine operation.

I also want to be able to investigate a slow table and change its configuration without waiting for someone else to give me access. Choosing a managed Cassandra service means working out how much of that control we keep and what help we’re getting in return.

Suppose we’re moving an application that stores customer activity in partitions organised by customer and day. The first query against the new service returns the expected rows and the driver connects without errors. We still need to check whether old activity expires as expected and whether we can recover a customer’s records after an accidental deletion. Then we need to price the workload we’re actually going to run.

We develop AxonOps so I’ll also explain how we approach self-hosted Cassandra operations as we work through the comparison. I’ve linked the provider documentation for the service details and checked those references on 16 September 2026.

Cassandra Hosting Options

Instaclustr and Azure Managed Instance run Apache Cassandra clusters with the provider handling agreed maintenance and administration. Amazon Keyspaces and Azure Cosmos DB expose Cassandra-compatible interfaces with different operational controls and application behaviour. Astra DB Serverless is powered by Cassandra but also applies its own service limits.

OptionDatabase and accessWhat to check first
Self-hosted Apache CassandraApache Cassandra on infrastructure you controlCapacity and maintenance responsibilities within your team
InstaclustrManaged Apache Cassandra with customer cloud account optionsSupported versions and the division of operational work
Azure Managed Instance for Apache CassandraManaged open source Cassandra with hybrid ring supportAzure connectivity and supported administration commands
DataStax Astra DB ServerlessCassandra-powered service with CQL access and service-managed configurationSchema limits and the operations available through your chosen API
Amazon KeyspacesAWS serverless service with a Cassandra-compatible interfaceQuery compatibility and request capacity requirements
Azure Cosmos DB for Apache CassandraCosmos DB through a Cassandra-compatible interfaceTTL behaviour and supported CQL operations

A successful driver connection tells us that the application can reach the database but leaves most of its behaviour untested.

Amazon Keyspaces

Amazon Keyspaces takes host and JVM administration out of the customer-facing service. You can connect through Cassandra drivers but importing a Cassandra schema won’t reproduce its storage configuration. AWS documents that compaction and compression settings are ignored along with other Cassandra-specific tuning options in its functional differences guide.

The same guide confirms support for conditional inserts and updates as well as deletes through lightweight transactions. An application that uses IF NOT EXISTS therefore has a supported feature to test rather than a reason to rule Keyspaces out immediately.

For our customer-activity example we’d also need to test when deleted records stop appearing in queries. AWS documents asynchronous range deletion from July 2026 where a successful response confirms acceptance while matching rows may still be present. The application needs to check completion if its next operation depends on those rows having gone.

AWS supports service-specific credentials and IAM-based authentication for applications connecting through Cassandra drivers. Existing Cassandra users and role definitions need to be mapped to the service’s access controls during migration.

Azure Cosmos DB for Apache Cassandra

Cosmos DB’s Cassandra interface runs on Cosmos DB and is separate from Azure Managed Instance. Microsoft’s supported-features reference documents CQL compatibility together with the exceptions that need application testing.

For example TTLs and custom timestamps supplied with USING apply at row level rather than per cell. If our activity application updates one field while relying on another field’s expiry time we need to test that exact update sequence before migrating.

Conditional writes are supported but lightweight transactions are unavailable for accounts with multiple-region writes enabled. The service also lacks logged batch support so applications using those batches need changes before migration. Check these constraints against the application’s queries and the regional configuration you plan to use.

Azure Managed Instance for Apache Cassandra

Azure Managed Instance runs open source Cassandra and supports adding managed datacentres to an existing Cassandra ring. That gives a team with an on-premises deployment a way to evaluate a hybrid migration through Cassandra replication. The service also supports configuration overrides and exporting metrics to Prometheus or Grafana.

Microsoft’s management operations guide covers repair and patching alongside backup and supported administration commands. Access extends beyond read-only nodetool commands but some operations carry explicit support restrictions.

The documented snapshot restore process starts with a support request so include the provider’s response time in your recovery rehearsal. Agree how the restored tables will be checked and who will reconnect the application once they’re ready.

Instaclustr

Instaclustr manages Apache Cassandra and offers deployments in a customer’s own cloud account. Its custom VPC documentation describes running a cluster inside a network the customer manages. That can be useful when the application already depends on private connectivity and an established cloud security setup.

The operating agreement still needs to cover configuration changes and upgrades alongside incident response. Ask which Cassandra versions are available for the proposed deployment and how a change gets approved during an incident. Running in your own account doesn’t by itself tell you which commands your engineers can execute.

Include the management service in the cost comparison and confirm what the quote covers. The amount of engineering work it takes over will depend on the service agreement and the work your team already automates.

DataStax Astra DB Serverless

DataStax describes Astra DB Serverless as powered by Apache Cassandra with the operational controls documented in its database limits. There is no JMX or nodetool access and the service controls cassandra.yaml while using Unified Compaction Strategy.

Astra supports Storage Attached Indexing with documented table limits of 75 columns and 10 indexes per table. A single column value can be up to 10 MB under the current limits. Check database-wide limits too when sizing a schema with many indexed tables.

The CQL guide explains that unsupported table properties can be ignored with a warning while the statement succeeds. Review those warnings during a schema import because a table created successfully may have different settings from the source.

The Data API uses LOCAL_QUORUM for reads and writes while CQL drivers have a wider choice of consistency levels subject to the documented restrictions. Test through the interface your application will use so the results reflect its configuration.

Aiven for Apache Cassandra

Aiven’s end-of-life notice lists 7 January 2026 as the end of its managed Cassandra service so it is no longer a hosting option for this comparison.

Include an exit plan when evaluating any provider and work out how to export the data while keeping track of writes during a move. You can test that process before a contract deadline makes those decisions urgent.

Application Compatibility

For the customer-activity application we’d want to know how queries behave as the history grows. Reading ten recent events is a useful first test but it won’t tell us how the service behaves for a customer with several years of records.

Use the actual driver version and production consistency settings when checking the following behaviours. Include the largest expected partitions and the update patterns used by the application.

Application operationTest to run
Read recent activityQuery a large partition and page through the complete result set
Expire old recordsUpdate individual fields and check which values expire at each TTL
Register a request onceRun concurrent conditional writes and verify retry behaviour
Search an indexed fieldRun the actual predicates against the intended index type
Delete customer activityConfirm when deletion completes and what concurrent readers can see
Apply a schema changeCheck accepted settings and capture any ignored-property warnings
Recover deleted recordsRestore to a separate environment and validate the application’s data

These tests can change the shortlist before we spend time negotiating prices. They also give us a repeatable way to check a future migration away from the selected service.

Costs

A monthly comparison needs to include the same application workload and recovery requirements on both sides. A bill for six virtual machines leaves out engineering time and backup storage while a serverless estimate based only on requests per second leaves out the size and cost of those requests.

OptionCapacity billing to evaluateAdditional costs to include
Self-hosted CassandraCompute and disks sized through workload testingNetwork transfer and backup storage alongside operations tooling and engineering time
InstaclustrA quote for the chosen cluster and management serviceConfirm which infrastructure and support charges are included
Azure Managed InstanceSelected VM sizes and disksAdditional backup copies and applicable network charges
Amazon KeyspacesOn-demand request units or provisioned throughputStorage and optional features such as point-in-time recovery
Astra DB ServerlessMetered requests or provisioned capacityStorage and transfer charges under the chosen plan
Cosmos DB for Apache CassandraRequest-unit capacity for the workloadStorage and the chosen regional deployment

Use the current Keyspaces pricing and Azure Managed Instance pricing for the selected region. For Cosmos DB the request-unit documentation explains how database operations consume capacity.

Astra also offers Provisioned Capacity Units on its Enterprise plan with capacity-based billing. Databases assigned to those groups aren’t billed for individual read and write request units so a comparison based exclusively on per-request pricing would miss an available option.

Request Volume

Suppose our application makes 10000 point reads and 5000 single-row writes every second for a 30-day month. For this example each read consumes one Keyspaces read request unit at LOCAL_QUORUM and each write consumes one write request unit. This assumes reads of up to 4 KB and writes of up to 1 KB as measured under AWS’s on-demand capacity rules.

OperationConstant rateVolume over 30 days
Point reads10000 per second25.92 billion read request units
Single-row writes5000 per second12.96 billion write request units

If the regional prices are quoted per million units the request bill is 25920 times the read-unit price plus 12960 times the write-unit price. This excludes storage and other charges and assumes a single-region deployment without extra operations or retries. Repeat the calculation using provisioned throughput pricing if that is an option for your workload.

Now change a read to return 10 KB and it needs three read units at LOCAL_QUORUM. The application’s request rate hasn’t changed but the read-unit total has tripled. Use the record sizes and query results from the application when estimating those units.

For self-hosting we’d measure how many nodes can sustain this workload at the required p99 latency while leaving capacity for repair and a failed node. Assigning it to six nodes without that test would give us a compute price without establishing whether the cluster could run the application.

Storage and Operations

Retaining 100 GiB of new logical data per day for 30 days gives us 3000 GiB before replication and compression at steady state. With replication factor three it becomes 9000 GiB of uncompressed replicated data before indexes and other storage overhead.

Disk sizing then needs room for compaction and recovery alongside the effect of the chosen compression settings. Compare that with the provider’s definition of billed storage rather than multiplying every managed storage estimate by three automatically.

Include backup retention and the cost of a restore rehearsal in both budgets. Then add the engineering time needed for the responsibilities your team keeps under each option. Self-hosting can offer substantial savings for a sustained workload but the saving needs to come from this complete comparison.

Control and Portability

Self-hosting gives us control over the Cassandra version and host configuration along with access to the SSTables. We can investigate a table’s read amplification and test a different compaction configuration on infrastructure we manage.

The Apache Cassandra 5.0 release brought SAI and vector search alongside trie-based storage improvements and Unified Compaction Strategy. Cassandra 5.0 also added JDK 17 support for running the database on a newer JVM. Check the version and JVM support in a managed service separately because a Cassandra-compatible endpoint may expose different features.

Keeping Apache Cassandra also lets us plan moves between cloud infrastructure and our own hardware. The transfer method still depends on version compatibility and access to backups or streaming. A migration needs a tested way to handle writes made during the move and a rollback procedure the team can actually run.

With a managed service I’d ask for that procedure during the evaluation. Find out whether the available export can preserve the timestamps and TTLs the application relies on and how long it takes with a realistic volume of data.

Running Cassandra with AxonOps

We built AxonOps because we wanted to keep that control while reducing the maintenance involved in running Cassandra. You retain the database and infrastructure while AxonOps brings monitoring and routine operations into one place.

AxonOps Cassandra monitoring collects metrics at five-second resolution alongside logs and operational events. The dashboards organise the metrics by their purpose so a large collection doesn’t become a screen full of unrelated charts.

For our activity application that means following a rise in query latency through the affected tables and checking SSTables per read alongside tombstone scans. We can compare the same measurements before and after a change with the logs and maintenance events from that period. Our Datadog Cassandra metric comparison covers the table-level detail that these investigations need.

Operational workHow AxonOps helps
Monitoring and diagnosisCassandra-specific dashboards with five-second metrics and correlated logs and events
RepairAdaptive regulation adjusts repair activity using Cassandra and host performance measurements
Backup and restoreScheduled backups and restore workflows with point-in-time recovery when the required archiving is configured
Routine maintenanceCentral scheduling and execution of supported operational tasks

Our article on adaptive Cassandra repair explains how live performance measurements guide repair activity. The backup and restore guide covers recovery options and supported storage targets.

Your team remains responsible for infrastructure capacity and recovery testing with AxonOps providing the tools to carry out that work. You can manage the clusters through one platform without maintaining a separate collection of monitoring and maintenance tools.

Choosing a Deployment

For an existing Cassandra application I’d start by evaluating self-hosting with AxonOps alongside the managed services that pass its compatibility tests. Run the same workload on each candidate and include a failed-node test or the provider’s supported failure exercise. Restore a backup and measure how long it takes to get a validated application running again.

A managed service can be a sensible choice when the work it takes over justifies the price and its controls meet the application’s needs. Self-hosting deserves the same practical evaluation because retaining access to Cassandra can make investigation and tuning easier while giving you more choice over infrastructure costs.

You can explore AxonOps in the demo sandbox or discuss your Cassandra deployment with us to see how it would fit the cluster you’re comparing.

All articles