Skip to content

Release Notes RonDB 25.10.19#

RonDB 25.10.19 is a release of the RonDB 25.10 series. It is a GA version of RonDB 25.10.

RonDB 25.10.19 is based on MySQL NDB Cluster 8.4.6 and RonDB 24.10.20.

RonDB 21.04 is a Long-Term Support version of RonDB that is no longer supported.

RonDB 22.10 is a Long-Term support version and will be maintained at least until 2026.

RonDB 24.10 is a Long-Term support version and will be maintained at least until 2027.

RonDB 25.10 is a Long-Term support version and will be maintained at least until 2028.

RonDB 26.02 is a Long-Term support version and will be maintained at least until 2028.

RonDB 25.10 is released as open source SW with binary tarballs for usage in Linux. It is developed on Linux and Mac OS X and using WSL 2 on Windows (Linux on Windows).

RonDB 25.10.2 and onwards is supported on Linux/x86_64 and Linux/ARM64.

RonDB 25.10.19 is also available as a Docker container for both x86_64 and ARM CPUs and one can also use RonDB through rondb-helm, a Helm chart to use RonDB in Kubernetes.

The other platforms are currently for development and testing. Mac OS X is a development platform and will continue to be so.

Description of RonDB#

RonDB is designed to be used in a managed cloud environment where the user only needs to specify the type of the virtual machine used by the various node types. RonDB has the features required to build a fully automated managed RonDB solution.

It is designed for applications requiring the combination of low latency, high availability, high throughput and scalable storage (LATS).

You can use RonDB in a Serverless version on app.hopsworks.ai. In this case Hopsworks manages the RonDB cluster and you can use it for your machine learning applications. You can use this version for free with certain quotas on the number of Feature Groups (tables) you are allowed to add and quotas on the memory usage. You can get started in a minute with this, no need to setup any database cluster and worry about its configuration, it is all taken care of.

You can use the open source version and use the binary tarball and set it up yourself.

You can use the open source version and build and set it up yourself.

These are the commands you can use to retrieve the binary tarball:

# Download x86_64 on Linux
wget https://repo.hops.works/master/rondb-25.10.19-linux-glibc2.28-x86_64.tar.gz
# Download ARM64 on Linux
wget https://repo.hops.works/master/rondb-25.10.19-linux-glibc2.28-arm64_v8.tar.gz

These versions are also available as Docker containers at hub.docker.com under hopsworks/rondb. See https://hub.docker.com/r/hopsworks/rondb/tags for an up to date list of which RonDB container images are available. These containers can be used to run RonDB in containers, in Hopsworks they are also used since version 4.0 as containers in the Hopsworks Kubernetes cluster.

The actions to build both x86_64 and ARM64 tarballs have now been fully automated.

Summary of changes in RonDB 25.10.19#

RonDB 25.10.19 is based on MySQL NDB Cluster 8.4.6 and RonDB 24.10.20.

RonDB 25.10.19 contains all changes since RonDB 25.10.9. It adds 11 improvements and fixes more than 35 bugs, many of them in the TTL feature.

New features in 25.10.19 include a restart barrier that makes rolling restarts across node groups safe, partition hash fanout tables, a daily active window and partition sharding for TTL purging, fine-grained sharing in the REST API server and better support for running the REST API server in Kubernetes.

Test environment#

RonDB uses four different ways of testing. MTR is a functional test framework built using SQL statements to test RonDB.

The Autotest framework is specifically designed to test RonDB using the NDB API. The Autotest is mainly focused on testing high availability features and performs thousands of restarts using error injection as part of a full test suite run.

Benchmark testing ensures that we maintain the throughput and latency that is unique to RonDB. The benchmark suites used are integrated into the RonDB binary tarball making it very straightforward to run benchmarks for RonDB.

Finally we also test RonDB in the Hopsworks environment where we perform both normal actions as well as many actions to manage the RonDB clusters.

RonDB has a number of MTR tests that are executed as part of the build process to improve the performance of RonDB.

MTR testing#

RonDB has a functional test suite using the MTR (MySQL Test Run) that executes more than 500 RonDB specific test programs. In addition there are thousands of test cases for the MySQL functionality. MTR is executed on both Mac OS X and Linux.

We also have a special mode of MTR testing where we can run with different versions of RonDB in the same cluster to verify our support of online software upgrade.

Autotest#

RonDB is very focused on high availability. This is tested using a test infrastructure we call Autotest. It contains also many hundreds of test variants that takes around 36 hours to execute the full set. One test run with Autotest uses a specific configuration of RonDB. We execute multiple such configurations varying the number of data nodes, the replication factor and the thread and memory setup.

An important part of this testing framework is that it uses error injection. This means that we can test exactly what will happen if we crash in very specific situations, if we run out of memory at specific points in the code and various ways of changing the timing by inserting small sleeps in critical paths of the code.

During one full test run of Autotest, RonDB nodes are restarted thousands of times in all sorts of critical situations.

Autotest currently runs on Linux with a large variety of CPUs, Linux distributions and even on Windows using WSL 2 with Ubuntu.

Benchmark testing#

We test RonDB using the Sysbench test suite, DBT2 (an open source variant of TPC-C), flexAsynch (an internal key-value benchmark), DBT3 (an open source variant of TPC-H) and finally YCSB (Yahoo Cloud Serving Benchmark). We test the REST API server using benchmark programs written in Go and Feature Store REST API server is benchmarked using a tool called Locust that can be used to test various application variants. This tool is heavily used to benchmark application scenarios required by Hopsworks customers.

Dydra, a community user of RonDB implements a graph database supporting SPARQL on top of RonDB using a Common Lisp NDB API. They also have a set of benchmarks to verify the performance of RonDB in their application.

The focus is on testing RonDBs LATS capabilities (low Latency, high Availability, high Throughput and scalable Storage).

Hopsworks testing#

Finally we also execute tests in Hopsworks to ensure that it works with HopsFS, the distributed file system built on top of RonDB, and HSFS, the Feature Store designed on top of RonDB, and together with all other use cases of RonDB in the Hopsworks framework.

Improvements#

RONDB-1096: Restart barrier for node restarts#

A restarting data node now waits in a new start phase 110 until all restarting data nodes have recovered, and graceful stops of other data nodes are refused meanwhile. This makes rolling restarts across node groups safe, e.g. in Kubernetes (new parameter RestartBarrierTimeout).

RONDB-1074: Partition hash fanout#

The table option COMMENT="NDB_TABLE=PARTITION_HASH=x:y:z" spreads the rows with the same first x primary key columns over z partitions, so that scans for one entity run in parallel on several partitions.

RONDB-1068, RONDB-1073: Improved CPU locking in automatic thread configuration#

Round Robin groups are sized by the query threads and LDM threads are spread over all Round Robin groups. MaxRRGroupSize can be set to 32, and the CPU binding topology is logged.

Daily active window for TTL purging#

TTL purging can be restricted to a daily UTC time window, set per cluster in mysql.ttl_purge_ctrl or per REST API server with TTLPurge.ActiveWindow.

TTL purge sharded by partition#

With several purge nodes, each purge node now purges a share of the partitions of every table instead of whole tables, so the purge rate of a table scales with the purge nodes.

Faster TTL expiry checks#

The expiry check reads the TTL column directly and compares integers, the current time is read once per scan batch, and the glibc time zone lock is avoided on the hot paths.

RONDB-1088: Fine-grained sharing in the REST API server#

API keys can be authorized for individual tables and columns shared with a Hopsworks project, not only for whole databases.

RONDB-1104: REST API server recovers from cluster outages#

The REST API server reconnects automatically, also when idle, after all data nodes were unavailable. /health reports unhealthy while no data node is reachable, and a cluster failure is reported as HTTP status 500 instead of 404.

RONDB-1134: REST API server improvements for Kubernetes#

A probe port (REST.ProbePort, default 4407) answers /ping and /health also when the data path is stalled. REST.MaxKeepaliveRequests and REST.IdleConnectionTimeoutS limit connection lifetimes so that a Kubernetes Service can rebalance clients, and REST.UploadPath sets the directory for large request bodies.

Data node start logging#

The node log describes the progress of a data node start with uniform [NODE-START] lines: the planned steps, the progress of each step and any stalls.

RONDB-1100: Removed the deprecated v1 REST server#

The source code of the v1 REST server, replaced by rest-server2 and no longer built, has been removed.

Bug Fixes#

RONDB-1067: Low load flag of block threads not maintained#

The low load flag used when distributing reads to query threads was not reset after it had been set once.

RONDB-1075: Index scan returning rows of another table#

A secondary index scan could return rows of another table after tables were dropped and recreated. API nodes now also invalidate their dictionary caches when a table is dropped.

RONDB-1092: Use-after-free in the NDB API dictionary cache#

A table object could be freed while still in use when the table was changed concurrently.

RONDB-1090: End of UNDO log not found after crash#

A crash during a multi-page UNDO log write could leave a hole before the last written pages, which made the search for the end of the UNDO log fail.

RONDB-1111: Lost TAKE_OVERTCCONF at master failure#

A master failure during a TC take over could duplicate or lose TAKE_OVERTCCONF signals and hang node failure handling.

RONDB-1112: Copy scan stall when the copy target fails#

A fragment copy during a node restart could stall when the copy target node failed.

RONDB-1114: LCP started before the LCP round#

LCP fragment checkpoints could start before the LCP round when an LCP participant failed during the start of the LCP.

Bug#36054459: TC node failure handling in progress when a node rejoins#

Backport of the MySQL fix for a node rejoining while TC node failure handling is still in progress.

Partial writes in the file system layer#

Short writes, e.g. after a signal or close to a full disk, were not retried and were reported as file append errors.

Data node abort due to wrong query thread ids#

Query thread ids were computed with a wrong base, which could abort data nodes with signal 6 (error 6000).

Data node crash on API_VERSION_REQ#

Data nodes crashed on an API_VERSION_REQ for a node id outside their configuration.

Node id locked after a failure during node start#

A data node that failed during its start could lock its node id permanently, and a FAIL_REP to a node that had not joined crashed instead of giving a graceful error.

Leaked LCP_SCANNED_BIT in the page map#

A leaked LCP_SCANNED_BIT could make the next LCP skip a live page. The leak is fixed, and leaked bits are now detected and healed.

ndb_mgmd exits during concurrent configuration changes#

An ndb_mgmd losing a race between concurrent configuration changes exited, and transient configuration check mismatches during a configuration change were treated as fatal.

RONDB-1134: Stale state in pooled NDB API objects#

Reused scan operations and transactions kept the aggregation program and error details of their previous use.

REST API server: API key table without expiry column#

The REST API server crashed when hopsworks.api_key had no expiry column, and the API key watcher did not survive ALTER TABLE of that table.

REST API server: schema errors invalidated the wrong table#

After a schema error, the REST API server invalidated another table than the one that was read.

REST API server: other fixes#

  • Signals delivered during startup now terminate the process.

  • Fixed a histogram bucket overflow in the RonSQL and index scan latency metrics.

  • Fixed JSON escaping of --print-config.

  • Fixed memory safety of early returns and retries in batch primary key reads.

TTL fixes#

  • Node restarts could fail when copying fragments of TTL tables (error 2303).

  • Restoring an LCP of a TTL table could fail with a row count mismatch (error 2352).

  • REDO log replay lost rows that expired while the node was down, and crashed on inserts over expired rows.

  • A node restart could silently lose rows when a write of a new key was ignored during the fragment copy.

  • Backup replicas now always receive the row id from the primary replica for TTL inserts and writes, preventing replica divergence (error 899). A mismatch is reported as error 1245.

  • Unique index constraints could be bypassed for TTL tables, also within one transaction, and a case-insensitive primary key collation gave false duplicate key errors.

  • Inserts over expired rows could leave stale unique index entries and leak BLOB parts.

  • An insert of a duplicate primary key in the same transaction became a silent upsert.

  • Data node crashes with deferred triggers, on the middle replica with NoOfReplicas of 3 or more, and with a TTL column in DYNAMIC format.

  • mysqld crashed on an integer overflow in the TTL table comment.

  • The replication applier could miss rows of TTL tables without primary key, causing replica divergence.

  • Foreign keys could be combined with TTL through ALTER TABLE, and TTL could be bound to an on-disk column through an inplace ALTER TABLE.

  • Renaming the TTL column left a stale binding in the table comment.

  • Deletes by key of TTL rows now only delete expired rows.

  • The TTL purge could stall on broken tables, miss out-of-band DDL, reset its rotation on reload and starve a copying ALTER TABLE.