Release Notes RonDB 25.10.9#
RonDB 25.10.9 is a release of the RonDB 25.10 series. It is a GA version of RonDB 25.10.
RonDB 25.10.9 is based on MySQL NDB Cluster 8.4.6 and RonDB 24.10.20.
RonDB 21.04 is a Long-Term Support version of RonDB that is no longer supported.
RonDB 22.10 is a Long-Term support version and will be maintained at least until 2026.
RonDB 24.10 is a Long-Term support version and will be maintained at least until 2027.
RonDB 25.10 is a Long-Term support version and will be maintained at least until 2028.
RonDB 25.10 is released as open source SW with binary tarballs for usage in Linux. It is developed on Linux and Mac OS X and using WSL 2 on Windows (Linux on Windows).
RonDB 25.10.2 and onwards is supported on Linux/x86_64 and Linux/ARM64.
RonDB 25.10.9 is also available as a Docker container for both x86_64 and ARM CPUs and one can also use RonDB through rondb-helm, a Helm chart to use RonDB in Kubernetes.
The other platforms are currently for development and testing. Mac OS X is a development platform and will continue to be so.
Description of RonDB#
RonDB is designed to be used in a managed cloud environment where the user only needs to specify the type of the virtual machine used by the various node types. RonDB has the features required to build a fully automated managed RonDB solution.
It is designed for applications requiring the combination of low latency, high availability, high throughput and scalable storage (LATS).
You can use RonDB in a Serverless version on app.hopsworks.ai. In this case Hopsworks manages the RonDB cluster and you can use it for your machine learning applications. You can use this version for free with certain quotas on the number of Feature Groups (tables) you are allowed to add and quotas on the memory usage. You can get started in a minute with this, no need to setup any database cluster and worry about its configuration, it is all taken care of.
You can use the managed version of RonDB available on hopsworks.ai. This sets up a RonDB cluster in your own AWS, Azure or GCP account using the Hopsworks managed software. This sets up a RonDB cluster provided a few details on the HW resources to use. These details can either be added through a web-based UI or using Terraform. The RonDB cluster is integrated with Hopsworks and can be used for both RonDB applications as well as for Hopsworks applications.
You can use the open source version and use the binary tarball and set it up yourself.
You can use the open source version and build and set it up yourself.
These are the commands you can use to retrieve the binary tarball:
# Download x86_64 on Linux
wget https://repo.hops.works/master/rondb-25.10.9-linux-glibc2.28-x86_64.tar.gz
# Download ARM64 on Linux
wget https://repo.hops.works/master/rondb-25.10.9-linux-glibc2.28-arm64_v8.tar.gz
These versions are also available as Docker containers at
hub.docker.com under hopsworks/rondb. See
https://hub.docker.com/r/hopsworks/rondb/tags for an up to date list
of which RonDB container images are available. These containers can be
used to run RonDB in containers, in Hopsworks they are also used since
version 4.0 as containers in the Hopsworks Kubernetes cluster.
The actions to build both x86_64 and ARM64 tarballs have now been
fully automated.
Summary of changes in RonDB 25.10#
RonDB 25.10.9 is based on MySQL NDB Cluster 8.4.6 and RonDB 24.10.20.
RonDB 25.10.9 adds 13 improvements on top of RonDB 24.10 and adds 82 improvements on top of MySQL NDB Cluster 8.4.6.
RonDB 25.10.9 has the same bug fixes as 24.10.20, but has the bug fixes from MySQL 8.4.5 and MySQL 8.4.6 as well. In addition it fixes 53 more bugs.
New features in 25.10 include support for more API nodes, a new index scan endpoint in the REST API server, a more complete implementation of RonSQL, performance improvements of long scans and significant improvements of restart times.
Test environment#
RonDB uses four different ways of testing. MTR is a functional test framework built using SQL statements to test RonDB.
The Autotest framework is specifically designed to test RonDB using the NDB API. The Autotest is mainly focused on testing high availability features and performs thousands of restarts using error injection as part of a full test suite run.
Benchmark testing ensures that we maintain the throughput and latency that is unique to RonDB. The benchmark suites used are integrated into the RonDB binary tarball making it very straightforward to run benchmarks for RonDB.
Finally we also test RonDB in the Hopsworks environment where we perform both normal actions as well as many actions to manage the RonDB clusters.
RonDB has a number of MTR tests that are executed as part of the build process to improve the performance of RonDB.
MTR testing#
RonDB has a functional test suite using the MTR (MySQL Test Run) that executes more than 500 RonDB specific test programs. In addition there are thousands of test cases for the MySQL functionality. MTR is executed on both Mac OS X and Linux.
We also have a special mode of MTR testing where we can run with different versions of RonDB in the same cluster to verify our support of online software upgrade.
Autotest#
RonDB is very focused on high availability. This is tested using a test infrastructure we call Autotest. It contains also many hundreds of test variants that takes around 36 hours to execute the full set. One test run with Autotest uses a specific configuration of RonDB. We execute multiple such configurations varying the number of data nodes, the replication factor and the thread and memory setup.
An important part of this testing framework is that it uses error injection. This means that we can test exactly what will happen if we crash in very specific situations, if we run out of memory at specific points in the code and various ways of changing the timing by inserting small sleeps in critical paths of the code.
During one full test run of Autotest, RonDB nodes are restarted thousands of times in all sorts of critical situations.
Autotest currently runs on Linux with a large variety of CPUs, Linux distributions and even on Windows using WSL 2 with Ubuntu.
Benchmark testing#
We test RonDB using the Sysbench test suite, DBT2 (an open source variant of TPC-C), flexAsynch (an internal key-value benchmark), DBT3 (an open source variant of TPC-H) and finally YCSB (Yahoo Cloud Serving Benchmark). We test the REST API server using benchmark programs written in Go and Feature Store REST API server is benchmarked using a tool called Locust that can be used to test various application variants. This tool is heavily used to benchmark application scenarios required by Hopsworks customers.
Dydra, a community user of RonDB implements a graph database supporting SPARQL on top of RonDB using a Common Lisp NDB API. They also have a set of benchmarks to verify the performance of RonDB in their application.
The focus is on testing RonDBs LATS capabilities (low Latency, high Availability, high Throughput and scalable Storage).
Hopsworks testing#
Finally we also execute tests in Hopsworks to ensure that it works with HopsFS, the distributed file system built on top of RonDB, and HSFS, the Feature Store designed on top of RonDB, and together with all other use cases of RonDB in the Hopsworks framework.
Improvements#
More API nodes#
Increase support of API nodes from 255 to 2039 (RONDB-914). This feature enables RonDB to handle 8x larger clusters.
Long scan improvement#
Support higher parallelism in scans (RONDB-934). In scans of millions of rows this improvement can speed up order by execution by a factor of 3-4 and significantly improve also other scans that execute for a long time.
Index rebuild during recovery improvement#
Single CPU used for index rebuild in CPU locking scenarios. Added capability to have exclusive CPUs used for IO activity, very useful with encrypted and compressed files.
Parallelise offline ordered index rebuild even more#
Parallelise offline ordered index rebuild even more. Very significant improvement of restart times with large amounts of ordered indexes.
Add many improvements to RonSQL#
Lots of improvements, new operators, more types supported, multi-column index scans and many fixes.
Add RDRS index scan feature with comprehensive data type support and observability#
The REST API can now handle scans with very low latency.
Many improvements to TTL implementation#
This includes bug fixes, REST endpoints for observability.
Add support for TIMESTAMP as primary key column in ClusterJ#
RONDB-1012: Improved MTR upgrade testing#
Add execv-based angel restart for upgrade/downgrade testing
The angel process uses fork()+real_main() to restart data nodes, which reuses the same in-memory binary. This prevents testing upgrade/downgrade scenarios where the restarted node should run a different binary version.
Add NDB_RESTART_WITH_EXEC env var support to angel.cpp that makes spawn/_process() use execv() instead of real_main(). MTR creates a symlink (ndbmtd_current -> ndbmtd/_v2) and starts nodes via it. Tests can update the symlink target before triggering ndb_mgm restarts, causing the angel to load the new binary from disk.
New test cases ndb_restart_upgrade and ndb_fk_upgrade_restart exercise this by starting with ndbmtd_v2 and restarting with ndbmtd. Both tests skip when ndbmtd_v2 is not available.
Fix angel to use execv with symlink path for upgrade testing
The angel process must be the new binary (ndbmtd) to have execv support. Previously, starting via the symlink meant the angel was ndbmtd_v2 (old binary without execv), so restarts always reused the same in-memory binary.
Now NDB_RESTART_WITH_EXEC contains the symlink path (not just "1"), and the angel uses that path in execv() when spawning children. MTR starts ndbmtd directly as the angel, while children are spawned through the symlink whose target can be changed between restarts.
Add per-test cnf for upgrade tests and skip incompatible options
Create separate .cnf files for ndb_restart_upgrade and ndb_fk_upgrade_restart with small node IDs that ndbmtd_v2 supports. The suite-level my.cnf keeps the original large node IDs for all non-upgrade tests.
Skip --ndb-log-timestamps when spawning via restart_exec_path since the child may be an older binary that does not recognise this option.
Update ndb_fk_upgrade_restart result file for upgrade echo messages
Remove mysqld_v2 and ndb_mgmd_v2 support, fix upgrade test skip logic
Only ndbmtd version mixing is supported via the angel/execv approach. Remove the counter-based mysqld_v2 and ndb_mgmd_v2 alternation code which never worked properly for in-test restarts.
Also require MTR_RONDB_V2 to be set (in addition to the ndbmtd_v2 binary existing) before setting HAVE_NDBMTD_V2, so upgrade tests correctly skip when the env var is not set.
Explicitly delete HAVE_NDBMTD_V2 when upgrade testing is not enabled
Prevents the env var from leaking from a parent process and causing upgrade tests to run when MTR_RONDB_V2 is not set or set to 0.
RONDB-1017: Add SET MaxDiskWriteSpeed command to MGM client#
Add a proper MGM client command to set MaxDiskWriteSpeed at runtime,
replacing the need to use the raw DUMP 100002 command. The new command
updates the persistent configuration, the runtime value on data nodes,
and the node’s local ConfigValues so ndbinfo reports the new value.
Syntax:
<id> SET MaxDiskWriteSpeed <value>
ALL SET MaxDiskWriteSpeed 100M
The implementation uses a semi-generic SetConfigParam signal that
carries a config key and Uint64 value, making it easy to extend for
additional parameters in the future. Adding a new parameter only
requires changes in Cmvmi.cpp (runtime dispatch) and
CommandInterpreter.cpp (client-side command).
Includes MTR test (ndb_config_set) and CLAUDE.md developer guide.
RONDB-1020 Implement ORDER BY execution in RonSQL ResultPrinter#
Add comprehensive EXPLAIN tests for ORDER BY.
RONDB-1022: Add preloading and NDB event-driven updates for API key and FS metadata caches#
Eliminate cold-start latency and keep caches consistent by preloading all entries at startup and watching NDB events for real-time updates.
API key cache:
-
Preload all keys at startup using parallel threads
-
NDB event watcher for
INSERT(load),UPDATE(refresh),DELETE(invalidate) -
Consolidated refresh thread validates cached keys against DB periodically
-
In-memory SHA256 verification using cached secret/salt (avoids DB round-trip)
Feature store metadata cache:
-
Preload all feature view metadata at startup using parallel threads
-
NDB event watcher for
INSERT(load) andDELETE(evict) -
Parallel preloading with configurable thread count (
preloadThreads)
Reconnect resilience:
-
On event watcher reconnect after subscription gap, preload picks up missed
INSERTs and refresh thread handles missedDELETEs/UPDATEs -
Per-instance UUID in NDB event names prevents collisions across RDRS instances
-
force_reconnect()test hook for end-to-end reconnect testing
Other fixes:
-
Fix
ref_count=0race inload_single_key(protect against eviction during async DB fetch) -
Extract
generate_event_uuid()to sharedndb_event_utils.hpp -
Defer cleanup in Go feature_store event test
RONDB-822: Add charset-aware group-by comparison, Blob/Text validation, and MySQL-compatible arithmetic fixes#
-
Add collation-aware
GBHashEntryCmpfor pushdown aggregation group-by comparison on both data node and API side, withmemcmpfast path when all group-by columns are non-charset-aware (all_binary_cmp) -
Reject Blob/Text columns in
NdbAggregator::GroupBy()since they have no valid comparison function (m_cmp = nullptr) -
Fix arithmetic operations to match MySQL behavior:
MODuses dividend’s sign only,INT_MIN64 * n(n != 1) returns overflow, subtraction and division setis_unsignedcorrectly -
Add comprehensive test program (
ndbapi_agg_test) with 18 test cases covering charsetGROUP BY, numericGROUP BY,NULLhandling, arithmetic edge cases, and Blob/Text rejection -
Add MTR integration test (
ndbapi_agg_test) followingndbapi_vss_testpattern
Fix rdrs2 ordered index scan to not require index keys in readColumns#
Switch SF_OrderBy to SF_OrderByFull so NDB injects index key columns
into the kernel-side result mask itself, instead of requiring callers to
list them in readColumns. The JSON projection is still driven by
rdrs2’s own read_columns vector, so the response shape is unchanged.
Pre-patch, ordered scans with a partial readColumns set hit NDB error
4341.
Adds orderby_full_test.go covering secondary composite and primary
indexes with zero / partial / full / read-all readColumns, ASC and
DESC, plus a multi-fragment merge-sort variant on big_tbl.
Side-fix in index_scan_helper.go: SQL ORDER BY now repeats the
direction per key column. SQL “ORDER BY a, b DESC” means
“a ASC, b DESC”, which disagrees with NDB’s SF_Descending across all
key columns.
Bug Fixes#
RONDB-981#
Crash when using FetchAllowed=2 with parallel sorted index scan.
RONDB-984: Single CPU used for index rebuild in CPU locking scenarios#
When using AutomaticThreadConfig=1 and NumCPUs=0 the main thread got locked to a single CPU. No specific locking was set for IO threads (also used as index rebuild threads). This meant that the IO threads inherited CPU locking from its parent thread (the main thread) which effectively limited the CPUs available for index rebuild and IO threads to the same CPU as the main thread.
The solution is to allow the IO threads to use all CPUs available to the node. This change only affects the scenario with AutomaticThreadConfig=1 and NumCPUs=0, setting NumCPUs means no CPU locking is used at all.
A new configuration parameter ExclusiveIoCPUs was added to make it possible to have exclusive CPUs for IO threads. It is limited to at most 8 CPUs and also limited to at most 25
The IO threads will use the same CPU locking also for index rebuilds. The fix of this bug can improve restart times by significant numbers with large data sets and many indexes.
RONDB-934: Fix upgrade information#
RONDB-985: Parallelise offline ordered index rebuild even more#
Prepare DBLQH for more parallelism in rebuild of ordered indexes More parallelisation of ordered index rebuild in DBTUP
Each DBLQH instance will send up to 4 parallel index builds to DBTUP. Each of those instances will handle up to 4 parallel index partition builds. Normally the number of LDM instances is half of the number of CPUs. Thus having each of those handling up to 4 instances means that we will have threads to handle up to twice the number of CPUs in the node. This should be sufficient to keep all CPUs busy without overloading them with too many queued threads.
RONDB-983: RonSQL: Add DATE_ADD, DATETIME, NOT, TIMESTAMP, XOR; Improve DATE, INTERVAL, WHERE#
-
Test case
ronsql_timestampforDATE,DATETIME,TIMESTAMPcomparisons in filters (table scan) -
Improve
WHEREexpression handling-
Add
simplify_ce()to simplify the expression, with a reasonable tendency towards conjunctive normal form. -
Allow comparisons with a left-hand side other than column name.
-
More stringent checking of string literals interpreted as dates.
-
-
Add filter functionality, i.e.
WHEREexpressions not executed as index ranges.-
XOR -
NOT -
Compare
DATEcolumn to a string literal orDATE_ADDexpression. (Previously onlyDATE_SUBwas supported.) -
Compare
TIMESTAMP/DATETIMEcolumn to string literal orDATE_ADD/DATE_SUBexpression. -
Allow nested
DATE_ADDandDATE_SUB. -
Add
INTERVALtypesMICROSECOND,SECOND,MINUTE,HOUR,WEEK,MONTH,QUARTERandYEAR. (Previously onlyDAYwas supported.) -
Data type and range checks when comparing an integer column to a constant.
-
Compare
VARCHARcolumn to string literal. -
Negative
INTERVALs.
-
-
Refactoring to support the above.
CAVEAT:
- No support for
TIMESTAMPtime zone handling. RonSQL will act as if server is in UTC.
-
Index scan completely rewritten.
-
Code for data type support is now shared with the filter code. This adds temporal and
VARCHARdata type support to index scans. -
Index choice algorithm completely rewritten. It is based only on metadata, which is expected to be enough for many use cases.
-
Use
NdbScanFilter::setSqlCmpSemantics() -
Add test cases combining
NULL,NOT,AND,OR,=,!=
-
Add test case
ronsql_hopsworks, a complex, thorough test case. The basic idea is to test all data types (integral, float, varchar, temporal) in all roles (selected column, aggregated column, index equality/range bound, filter equality/range condition,GROUP BYcolumn). So far only the 10 integer types are tested. -
Improve
EXPLAINformat:-
Indent
WHEREconditions. -
Distinguish index bounds from filters.
-
Use boxdraw characters for tree structures.
-
-
Rewrite error and schema handling.
-
Rework exception types to more accurately predict whether schema unload and retry will work.
-
When necessary, reload schema object and compare id/version numbers to determine whether the error is retryable.
-
Improve use of ndb error information to determine whether the error is retryable.
-
Use
requirerather thanassert. -
Rename
assert_*torequire_*. -
Adjust messages.
-
Increase RonSQL query max attempts from 2 to 3 in
rdrs2andronsql_cli. -
Add error handling information to error stream.
-
-
Improve
ronsql_compare.inc-
Shorten diff separator line
-
Add option to suppress
ronsql_clicall, meaning only RDRS will be tested. Use this as necessary to speed up the test suite.
-
-
Improve debug messages
-
Add
DEB_TRACE()inNdbScanOperation.cppandRonSQLPreparer.cpp. -
Add
DBGmacro to track some values inRonSQLPreparer.cpp.
-
-
Check
NdbTransaction::getNdbIndexScanOperation()return value forNULL.
RONDB-983: kernel/ndbapi: Fixes for Pushdown Aggregation#
AggInterpreter (kernel):
-
Fix and refactor reading
DECIMALcolumns:-
Fix instruction pointer bug when reading a
NULLvalue. -
Fix incorrect
decimal.lenvalue. -
Fix
decimal_buf_type.
-
-
Check that program does not end in the middle of an instruction, to avoid reading beyond the
prog_buffer size.
storage/ndb/src/kernel/blocks/dbtup/AggInterpreter.hpp:
NdbAggregator (ndbapi):
-
Fix
Column::data_doublereturn type. -
In
ResultRecord::FetchGroupbyColumn, checkgetColumnreturn value.
NdbScanOperation (ndbapi):
-
Check table version in
setAggregationCode(). -
In
DoAggregation(), fix handling ofnextResultreturn value.
RONDB-983: RonSQL: Improve FLOAT, DOUBLE, DECIMAL (UNSIGNED), VARCHAR, DATE, DATETIME and TIMESTAMP#
ronsql_hopsworks test case:
-
Thoroughly test
FLOAT,DOUBLE,DECIMAL,DECIMAL UNSIGNED,DATE,DATETIME,TIMESTAMPandVARCHAR. -
Do not use
ORDER BY. -
Run
EXPLAINafter query. -
Test filter choice algorithm (goodness).
ResultPrinter:
-
Add support for printing
GROUP BYcolumns (non-aggregatedSELECTexpressions) of typesFLOAT,DOUBLE,DECIMAL,DECIMAL UNSIGNED,VARCHAR,DATETIMEandTIMESTAMP. -
Fix printing for
GROUP BYcolumns of typeDATE, and all aggregates.
Lexer:
- Store both double value and original string for float literals. This allows later convertion to decimal without loss of precision.
Parser:
-
Allow float literals in conditional expressions.
-
Remove negation simplification (moved to
RonSQLPreparer).
RonSQLPreparer:
-
Improve filter choice algorithm to prefer
PRIMARYindex. -
Add support for index range and comparison filters for columns of type
FLOAT,DOUBLE,DECIMALandDECIMAL UNSIGNED.- Change
raw_value::lenfromUint32tosize_tfor consistency withdecimal_bin_size()return type.
- Change
-
Add simplification of negation for integer and float literals.
-
Add support for printing float literals and intermediate temporal values in
EXPLAIN. -
Fix
ResultPrinterto never use scientific notation. -
Fix schema handling
-
Add
m_indexes. -
Fix
load(),generate_scan_config_candidates(),execute()andunload_schema().
-
-
Add debug prints.
-
Formatting.
RONDB-991: The RonDB REST API pk-read endpoint was returning null values for columns defined with NOT NULL constraints#
Root Cause: For NOT NULL columns, NDB’s default record can have null bit storage that overlaps with actual data storage. The code was checking the null bits without verifying if the column was actually nullable, causing it to misinterpret data bytes as null indicators. Added a check for col->getNullable()to only check null bits for columns that are actually nullable.
Fix TTL bugs in ALTER TABLE handling, scan options, and table comment parsing#
-
AlterTable.hpp:getTTLColFlag()/setTTLColFlag()usedTTL_SEC_SHIFTinstead ofTTL_COL_SHIFT, causing incorrect flag read/write -
DbtupMeta.cpp:handleAlterTablePrepare()checkedgetTTLSecFlag()twice instead ofgetTTLSecFlag() || getTTLColFlag() -
ha_ndbcluster.cc:full_table_scan()overwroteoptionsPresentwith=instead of|=forSO_TTL_IGNORE, losing previously set options -
ha_ndbcluster.cc:check_if_supported_inplace_alter()checkedmod_ttl != nullptrinstead ofmod_ttl->m_found
Set TTL ignore on copy-fragment scan during node restart#
Copy-fragment scans must see all rows including expired ones, otherwise expired rows would be lost during node restart. Set m_ttl_ignore = 1 on the scan operation in execCOPY_FRAGREQ.
Add REST API for dynamic TTL purge control and observability#
Introduce /0.1.0/ttl-purge REST endpoints to dynamically control TTL
purging without restarting RDRS:
-
GETconfig/status/metrics/tablesendpoints for monitoring -
PUTconfigendpoint to enable/disable purging and adjust batch size and sleep interval at runtime -
Per-table metrics tracking (rows purged, batches, errors)
-
Thread-safe design with
shared_mutex(read-once-per-round pattern)
Also add test scripts for TTL purge API and TTL test suite with environment variable overrides for portable test configuration.
Fix TTL purge bugs: potential infinite retry loop, unchecked NDB API calls, and partition contention#
-
Fix infinite retry loop where pre-transaction errors reset
trx_failure_timesper-table, preventing escalation. Addpre_trx_failurescounter at do-while scope that skips failing tables and triggers full restart after threshold. -
Fix
bool-typedcheckvariable inGetShardthat made scan error detection (check == -1) always false. Change tointto matchnextResult()return type. -
Add return value checks for
op->equal()inGetPurgeWindowandupdateTuple()/equal()/setValue()inUpdateLease. Remove unused variable. -
Offset starting partition by
nodeIdin no-sharding mode so multiple RDRS nodes scan different partitions, reducing lock contention.
Add backup-restore TTL test cases and fix metadata lock issues in test suite#
New test cases for the backup_restore category:
-
46c: Verify TTL metadata is preserved after restore (insert-after-restore still expires)
-
46d: Verify rows with
NULLTTL column survive backup-restore -
46f: Verify
ALTER TTL(disable, re-enable, change duration) works after restore -
46g: Verify disk columns + TTL survive backup-restore cycle
-
46h: Verify multiple tables with different TTL configs retain independent TTL behavior after restore
Fix existing backup-restore tests (46a, 46b) to:
-
Kill stale connections to
ttl_backupbeforeDROP DATABASEto prevent metadata lock hangs from previous failed runs -
Stop and rebuild replication in setup/teardown of every backup-restore case so the full category can run sequentially without breaking replication state
RDRS: Enable index_scan in MTR and fix potential shutdown hang#
index_scan test: Improve Test_SchemaVersionChangeConcurrent to tolerate transient errors during schema changes and track success/failure stats separately. Enable test in MTR.
RONDB-1006: Fix -Warray-bounds warnings for NodeFailRep cast in ClusterMgr#
The compiler warns that casting NdbApiSignal’s theData[] to NodeFailRep* creates an object whose full size (including the large NodeBitmask union) exceeds the signal buffer bounds.
Replace CAST_PTR(NodeFailRep, ...) with direct Uint32* access using named index constants (FailNoIndex, MasterNodeIdIndex, NoOfNodesIndex) added to NodeFailRep. Only the 3 signal header words are accessed; the bitmask is already sent separately via LinearSectionPtr.
Fixes warnings in threadMain, reportDisconnected, and execNODE_FAILREP.
RONDB-1007: Fix GCC 12 -Wstringop-overflow false positive in testClusterJ#
GCC 12 over-inlines the std::string concatenation chain in Paths::libBuildDir() and loses track of the SSO/heap allocation boundary, producing spurious -Wstringop-overflow and -Wstringop-overread warnings.
Define libBuildDir() out-of-line to break the inlining chain.
RONDB-1008: Suppress GCC 12 false positive warnings in vendored simdjson/jsoncpp/drogon#
GCC 12 produces spurious -Wmaybe-uninitialized warnings from its own
AVX512 intrinsic headers (the __m128i __Y = __Y self-initialization
pattern in avx512fintrin.h/emmintrin.h) when compiling simdjson’s
icelake backend.
GCC 12 also produces a spurious -Wstringop-overflow warning in
jsoncpp’s unindent() due to over-inlining std::string::resize().
Suppress at the CMake level since these are vendored third-party sources that should not be modified.
RONDB-1009: Replace strncpy with snprintf in RDRS status helpers#
GCC 12 warns about strncpy potentially truncating when the copy length
equals the source length (-Wstringop-truncation). The code was correct
(null-termination was done manually), but snprintf handles both
copying and null-termination in one call and avoids the warning.
RONDB-1008: Suppress GCC 12 AVX512 false positive during LTO link of rdrs2#
The -Wno-maybe-uninitialized flag on the simdjson compile target does
not carry through to the LTO link phase of rdrs2, where GCC
re-analyzes the inlined AVX512 intrinsics and re-emits the same false
positive about __Y in avx512fintrin.h/emmintrin.h.
Add target_link_options with -Wno-maybe-uninitialized (GCC-only) to
the rdrs2 executable to suppress the LTO-phase warning.
RONDB-1008: Suppress GCC 12 -Wstringop-overflow in drogon test LTO linking#
GCC 12 over-inlines std::string concatenation in drogon’s test macros
(stringifyFuncCall in drogon_test.h) and produces spurious
-Wstringop-overflow warnings during the LTO link phase.
Add target_link_options with -Wno-stringop-overflow (GCC-only) to
the drogon test executables.
RONDB-1008: Also suppress -Wmaybe-uninitialized in drogon test LTO linking#
The drogon test executables link simdjson (via drogon) and hit the same
GCC 12 AVX512 intrinsic false positive during LTO as rdrs2.
RONDB-1008: Use add_link_options for drogon GCC 12 LTO warning suppression#
The per-target approach only covered 3 test executables, but there are
60+ drogon/trantor executables (tests, examples, tools) that all link
simdjson and hit the same GCC 12 false positives during LTO. Use
add_link_options before add_subdirectory to cover all targets in the
drogon scope.
RONDB-1008: Suppress GCC 12 -Wrestrict false positive in drogon HttpController#
GCC 12 over-inlines std::string::replace in drogon’s
HttpController.h template code and produces a spurious -Wrestrict
warning about memcpy overlap with impossible sizes (near SIZE_MAX).
Add -Wno-restrict to drogon PUBLIC compile options since the warning
originates from a header included by downstream targets.
RONDB-1008: Move -Wno-maybe-uninitialized link suppression to rest-server2 top level#
The simdjson AVX512 LTO false positive was only suppressed for rdrs2
and drogon-scoped targets, but api_key_test and feature_store_test
also link simdjson transitively through rdrs2_lib. Move the
add_link_options to the rest-server2 top-level CMakeLists.txt so it
covers all targets in the tree, and remove the now-redundant per-target
and per-directory suppressions.
RONDB-1008: Also suppress -Wstringop-overflow at rest-server2 top level#
The jsoncpp -Wstringop-overflow false positive from std::string
over-inlining reappears during LTO linking of targets outside the drogon
directory scope (rdrs2, test executables). Consolidate with the
existing -Wno-maybe-uninitialized at the rest-server2 top level and
remove the now-redundant drogon-level add_link_options.
RONDB-1010: Silence CMake CMP0075 policy warning from drogon#
Set CMP0075 to NEW before adding the drogon subdirectory so that check_include_file_cxx() honors CMAKE_REQUIRED_LIBRARIES without emitting a developer warning.
RONDB-1010: Silence CMake CMP0075 policy warning from base64#
Set CMP0075 to NEW so check_include_file() honors CMAKE_REQUIRED_LIBRARIES without emitting a developer warning.
RONDB-1011: Make -Wno-restrict GCC-only in drogon compile options#
RONDB-1011: Use portable printf format attribute in ClusterMgr.cpp#
Replace gnu_printf with printf in the format attribute of error_printer. No call sites use GNU extensions so the check is equivalent, and Clang does not recognize gnu_printf.
Clang does not have -Wrestrict and warns about the unknown option. Use a generator expression to apply it only on GCC.
RONDB-1008: Make -Wno-maybe-uninitialized GCC-only in simdjson compile options#
Clang does not have -Wmaybe-uninitialized and warns about the unknown option. Use a generator expression to apply it only on GCC.
RONDB-1008: Make -Wno-stringop-overflow GCC-only in Jsoncpp_lib compile options#
Clang does not have -Wstringop-overflow and warns about the unknown option. Use a generator expression to apply it only on GCC.
RONDB-1008: Suppress Clang -Wunneeded-internal-declaration in jsoncpp#
The static inline getDecimalPoint() in jsoncpp is unused, triggering a Clang-only warning. Guard with AppleClang since GCC does not have this flag.
RONDB-1008: Suppress -Wmissing-field-initializers in base64-bin#
The vendored base64 CLI uses NULL as a struct option sentinel which triggers -Wmissing-field-initializers on Clang.
RONDB-1008: Suppress -Wsign-compare in flex-generated RonSQLLexer#
The generated lexer code compares signed and unsigned integers in flex boilerplate. Suppress only on the generated file.
RONDB-1008: Fix undefined reinterpret_cast in spj_performance_test#
The resultPtrs array was typed as const Row ** but only used with the setResultRowRef API which takes const char *&. Change the array type to const char ** to match the API and remove the UB reinterpret_cast.
RONDB-1008: Suppress GCC 12 -Wrestrict false positive in merge_large_tests-t#
GCC 12 emits a false -Wrestrict from std::string::replace inlining in googletest’s CanonicalizeForStdLibVersioning (gtest-type-util.h).
RONDB-1008: Apply mysqld Boost warning suppression unconditionally#
Upstream MySQL guards the -Wno-error=maybe-uninitialized and -Wno-error=uninitialized link options behind WITH_LTO, but GCC 12 also triggers these false positives without LTO through deep inlining of Boost geometry/spirit headers. Remove the LTO guard.
RONDB-TTL: Fix AlterTabOperation pool leak for TTL-only ALTER TABLE#
TTL-only ALTER TABLE operations
(e.g. ALTER TABLE t COMMENT="NDB_TABLE=TTL=5@col_b") seized an
AlterTabOperation record in handleAlterTablePrepare() but never
released it in handleAlterTableCommit() or handleAlterTableAbort().
Each leaked record permanently consumed one slot from the 16-entry pool
per DBTUP instance, eventually crashing all data nodes with “Pointer too
large” (error 2306) at DbtupMeta.cpp:1100.
The fix has three parts:
-
handleAlterTableCommit-- TTL-only path: addreleaseAlterTabOpRec()after applying TTL values. Guarded by!getAddAttrFlagsince the combined AddAttr+TTL case shares theAlterTabOpwith AddAttr which already releases it. -
handleAlterTableCommit-- combined AddAttr+TTL path: apply TTL values inside the AddAttr block beforereleaseAlterTabOpRec(), instead of reading from the already-freed record in the separate TTL block (pre-existing use-after-free). -
handleAlterTableAbort: add else-if branch to release theAlterTabOpfor TTL-only aborts. No descriptor release needed since TTL-only AlterTabOps havenullptrdescriptors.
RONDB-1016: Fix ronsql_overflow crash by increasing Drogon thread stack size#
Drogon worker threads used std::thread with OS default stack size ( 512KB on macOS), which overflows when RonSQL processes deeply nested queries ( 9987 levels of recursion in AST traversal). Switch EventLoopThread from std::thread to pthread_create with configurable stack size, set to 8MB in RDRS2.
Add optional LIMIT clause to RonSQL queries#
RONDB-1021: Added rdrs_http_connection_count metric#
RONDB-1030: Disable NDB event merging for feature_view events#
With mergeEvents(true), INSERT + DELETE events for the same feature_view row within the same GCI epoch ( 2s) cancel each other out. This causes DELETE events to be silently lost, leaving stale entries in the metadata cache.
Feature view changes are rare schema operations, so the extra event volume from disabling merging is negligible.
RONDB-1030: Remove time-based eviction of valid FS metadata cache entries#
The cache_entry_updater thread was evicting valid entries after 30 minutes of inactivity (cacheUnusedEntriesEvictionMS). This defeats the purpose of preloading — evicted entries force the next request through the slow lazy-load path, reintroducing the tail latency that preload eliminates.
Valid entries should only leave the cache on DELETE events (feature view deleted) or server shutdown. This matches the API key cache behavior.
Remove the cacheUnusedEntriesEvictionMS config parameter and simplify cache_entry_updater to only drain entries during shutdown.
RONDB-1031: Fix capitalizeMember to handle non-ASCII and non-letter Avro field names#
RONDB-1040: API_FAILREQ should be ignored if failed node is not part of nodes configuration#
RONDB-1039: Fix off by one errors in upgrade checks#
RONDB-1042: Update go libs for go lang version 1.26.1#
Fix rdrs2 scan heap corruption on composite secondary index over multi-col PK#
CompileIndexRanges in rdrs_dal.cpp iterated the index record’s
attrIds in ascending base-table attrId order via
getFirstAttrId/getNextAttrId, but the user-supplied range values in
lower/upper.values are in user-specified index-column order. For a
composite secondary index on a table whose PK columns are interleaved
with the index columns in attrId space
(e.g. PK=(client_id, occurred_at, event_id), secondary
=(customer_number, occurred_at); cn attrId=3, ts attrId=1), the
two orders disagree and each value is memcpy’d at the wrong column’s
offset. A VARCHAR binary landing in a BIGINT slot spills past
m_row_size and corrupts adjacent heap metadata in m_scan_buffer;
glibc notices later at NdbScanOperation::close() and aborts rdrs2
with “corrupted size vs. prev_size”.
Derive curr_attrId from the user-supplied column
(node.col->getAttrId()) at each iteration instead of walking the
record’s attrIds. Symmetric change in both lower and upper loops.
Add MTR regression test
mysql-test/suite/rdrs2/t/rdrs2_scan_composite_index.test using the
minimal 3-row schema that triggers the attrId-order mismatch, with
both DESC and ASC ordered scans. Pre-fix it crashes rdrs2;
post-fix both return the expected rows. Full Go scan suite
(mysql-test/suite/rdrs2-golang/rdrs2-golang_index_scan) also green.
Also fix TTLPurger VARBINARY encoding of “API_OK”
(ttl_purge.cpp): setValue("message", ...) on the VARBINARY(255)
column in mysql.ndb_schema_result reads the first byte as a length
prefix, treating ’A’ (0x41=65) as length and memcpy’ing 1+65
bytes from a 7-byte literal. Junk on the wire in a normal build; under
ASan a fatal global-buffer-overflow that blocked every ASan MTR run and
was caught while regression-testing the scan fix. Framed as a proper
VARBINARY buffer: 1-byte length (6) + “API_OK”.
Bugfix ttl read backup blob 25.10#
Allow NdbBlob lock-upgrade scans to route to backup replica.
This task is a simple one-off conversion, so task tracking isn’t needed here.
DbtcMain.cpp:16438 ndbrequire was too strict for
TTL+BLOB+READ_BACKUP/ FULLY_REPLICATED scans:
NdbBlob::atPrepareNdbRecord upgrades the user’s LM_CommittedRead to
LM_Read for blob-read atomicity and stamps rcb=1. The outer if
accepts the scan (rcb=1, op_count=0), but the inner check
"!ttl_table || (!lockmode && !holdlock)" then fired on the now-set
holdlock and killed the data node with error 2341.
Allow rcb=1 through the assertion; mirrors the non-TTL policy already
documented at sendDihGetNodeReq (“safe to read from backup replicas”).
“Read what you locked, even if expired” is preserved because the outer
if still requires op_count==0, so no prior in-txn locks exist when
the replica branch fires.
* ndb_ttl: regression test for TTL+BLOB+READ_BACKUP scan crash
Autocommit SELECT of a BLOB column on a TTL table with
READ_BACKUP=1 or FULLY_REPLICATED=1 used to kill the data node
picked for the scan with ndbrequire at DbtcMain.cpp
sendDihGetNodesLab. This test runs both table shapes and expects the
two inserted rows back.
RONDB-1058: Prefer ndb_mgmd as arbitrator during restart#
Adds a new DB-level config ArbitrationRankWait (ms, default 60000) so
QMGR waits for a rank-1 candidate (typically ndb_mgmd) before falling
back to a rank-2 SQL node when electing an arbitrator. If a rank-1
candidate joins later while a rank-2 fallback is active, QMGR demotes
the rank-2 arbitrator and re-elects so the mgmd takes over.
The ndb_mgmd startup path now keeps the MGM service reachable for
data-node bootstrap while rejecting ordinary MGM client commands until
arbitration has settled. This lets data nodes allocate node ids, fetch
config, set ports, and establish transporter connections without making
readiness probes such as ndb_mgm show succeed too early. When a
connected data node lacks support for the new behaviour, the startup
gate is capped at 2 s to avoid penalising mixed-version upgrades.
MgmtSrvr::wait_until_arbitrator() also folds a cold-start check into
the wait loop. If no data node has connected within the cold-start
window, the gate returns early so a freshly-started cluster does not
idle waiting for an arbitrator role it cannot yet be assigned.
TransporterFacade gains the support predicates for connected DB nodes
and unsupported DB-node versions.
Move the mgmd active-arbitrator marker to the point where the arbitrator
thread has entered started state and sent ARBIT_STARTCONF back to
QMGR. The code documents why ARBIT_STARTREQ implies all current data
nodes have completed the PREP2 agreement on arbitrator node and ticket,
including the two-data-node president/non-president failure ordering.
Report rank-2 arbitrator demotion as a state transition via
reportArbitEvent so it lands in each data node’s local ndbd.log, not
only the cluster log. Adds ArbitCode::ApiDemoted and a matching
getTextArbitState line. Also mirrors the rank-2 fallback infoEvent
with g_eventLogger->info so fallback context is visible locally while
mgmd is down.
Adds MTR coverage under suite/ndb: ndb_arbitration_rank_wait
verifies the full handoff trail (mgmd \rightarrow rank-2 \rightarrow
mgmd), ndb_rolling_restart_mgmd_first walks a rolling restart starting
with mgmd, and ndb_arbitration_president_failover stops mgmd, waits
for rank-2 takeover, kills the president data-node child, restarts mgmd,
and verifies data nodes can be cycled afterward. The warm-restart tests
use restart:–config-change to bypass MTR’s
ndb_mgmd_wait_started --no-contact path.
Tighten rdrs2 scan input validation#
Reject limit<0, bound.values longer than the index key,
varchar/longvarchar CMP values with non-string JSON kind. Flip
filter depth check to >= so MAX_FILTER_DEPTH=32 caps at exactly 32
nesting levels.
RONDB-1063: Fix crash in large initial start due to ACTIVATE/DEACTIVATE#
ACTIVATE_REQ is sent before starting the data node in rondb-helm. To
stop the node we deactivate it. This means that when we start a node
while another node has started its start the ACTIVATE_REQ can arrive
even before we have had time go through Phase 1.
The code ensures that no one can connect (execCONNECT_REP) before
start phase 1 have executed. This phase sets my node phase (state
variable) to ZSTARTING. If ACTIVATE_REQ arrived before this we would
open communication to the activated node. This led sometimes to the
problem that the activated node connected before we came to phase 1.
The fix is fairly straightforward, avoid sending OPEN_COMORD if I
haven’t reached the phase 1 yet. This means that OPEN_COMORD will be
sent from phase 1 start for all data nodes. QMGR expected this call to
be the first OPEN_COMORD call. When it wasn’t we even executed
OPEN_COMORD in the wrong branch. So the fix ensured that the phase 1
OPEN_COMORD was restored as the first call to OPEN_COMORD.
We protect both calls to OPEN_COMORD in this manner although only the
ACTIVATE_REQ is the one we can reach in practice.
Added test case for it, it was a bit difficult to reproduce, we had to
fix the handling of OPEN_COMORD to reproduce the issue. Also we could
not follow the exact mechanism from rondb-helm since MTR fails when we
use deactivate. Instead we added a DUMP command to set the node as
deactivated.
NdbAggregator::ProcessRes: fix wrong aggregate-result offset for new groups#
When inserting a brand new group-by entry into gb_map_, agg_res_ptr
was computed as agg_rec + agg_res_len instead of
agg_rec + gb_cols_len. The record layout is
[gb_cols_len bytes of group-by data | agg_res_len bytes of aggregate results],
so the aggregate area starts at gb_cols_len. The “found existing
entry” branch already uses iter->second.ptr (set to
agg_rec + gb_cols_len at insert time); only the new-entry path had the
wrong offset, and writes to agg_res_ptr[i] spilled past the allocation
whenever agg_res_len > gb_cols_len, corrupting the heap.
Also reorder the per-item type-match assert in the merge loop so it only
fires when both res[i] and agg_res_ptr[i] are non-null. The previous
assert allowed agg_res_ptr[i].is_null but not res[i].is_null, so it
tripped spuriously whenever an incoming null arrived for a non-null
accumulator.
TTL purge: separate round and table error paths
Default-initialize TTLInfo fields and per-round scratch buffers so
exceptional paths do not read stale stack or default-constructed values.
Split PurgeWorkerJob error handling into round_err for pre-table
failures and table_err for failures while processing a valid
local_ttl_cache iterator. This prevents cleanup or no-TTL-table paths
from jumping into the table loop handler and continuing into ++iter
with no valid iterator.