Problem
ControlConnectionTests.Integration_Cassandra_TopologyChange intermittently fails after bootstrapping node 2. The test expects queries to use two hosts, but sometimes observes only node 1:
Expected: expected_nodes.size()
Which is: 2
To be equal to: hosts.size()
Which is: 1
Recent failure:
https://github.com/scylladb/cpp-rs-driver/actions/runs/36589927177/job/109483566286
The identical assertion failed on master at commit 6a4b16284469cbf001c8735e70b4c628cc2be496, before PR #509:
https://github.com/scylladb/cpp-rs-driver/actions/runs/34943697382/job/104300208405
Cause
The test waits for the Node added to cluster log and then immediately makes only expected_nodes.size() + 2 requests:
|
void check_hosts(Session session, const std::set<unsigned short>& expected_nodes) { |
|
// Execute multiple requests and store the hosts used |
|
std::set<std::string> hosts; |
|
for (size_t i = 0; i < (expected_nodes.size() + 2); ++i) { |
|
Statement statement("SELECT * FROM " + system_schema_keyspaces_); |
|
Result result = session.execute(statement, false); |
|
if (result.error_code() == CASS_OK) { |
|
std::string host = result.host(); |
|
if (!host.empty()) hosts.insert(host); |
|
} else { |
|
TEST_LOG_ERROR("Failed to query host:" << result.error_message() << "[" |
|
<< result.error_code() << "]"); |
|
} |
|
} |
|
|
|
// Validate the hosts used during request execution and the expected |
|
ASSERT_EQ(expected_nodes.size(), hosts.size()); |
|
for (std::set<unsigned short>::const_iterator it = expected_nodes.begin(); |
|
it != expected_nodes.end(); ++it) { |
|
std::stringstream node_ip_address; |
|
node_ip_address << ccm_->get_ip_prefix() << *it; |
|
ASSERT_GT(hosts.count(node_ip_address.str()), 0u); |
|
} |
|
} |
|
CASSANDRA_INTEGRATION_TEST_F(ControlConnectionTests, TopologyChange) { |
|
CHECK_FAILURE; |
|
is_test_chaotic_ = true; // Destroy the cluster after the test completes |
|
|
|
/* |
|
* Create a new session connection using the round robin load balancing policy |
|
* to ensure all nodes can be accessed during request execution |
|
*/ |
|
Cluster cluster = default_cluster().with_load_balance_round_robin(); |
|
Session session = cluster.connect(); |
|
|
|
// Bootstrap a second node and ensure all hosts are actively used |
|
// Match the tail of "Node added to cluster: <host_id> - <address>" since the |
|
// host_id is not known ahead of time. |
|
logger_.add_critera("- " + (ccm_->get_ip_prefix() + "2") + ":9042"); |
|
EXPECT_EQ(2u, ccm_->bootstrap_node()); // Triggers a `NEW_NODE` event |
|
EXPECT_TRUE(wait_for_logger(1)); |
|
std::set<unsigned short> expected_nodes; |
|
expected_nodes.insert(1); |
|
expected_nodes.insert(2); |
|
check_hosts(session, expected_nodes); |
|
|
|
/* |
|
* Decommission the bootstrapped node and ensure only the first node is |
|
* actively used |
|
*/ |
|
decommission_node(2); // Triggers a `REMOVE_NODE` event |
|
expected_nodes.erase(2); |
|
check_hosts(session, expected_nodes); |
The Rust driver emits the node-added log while constructing refreshed cluster state, before it finishes initializing all pools and publishes that state for query routing. The log therefore confirms discovery, not readiness for requests. Immediate queries can still use the old one-node state.
Expected behavior
The integration test should tolerate asynchronous topology publication while still failing deterministically when the driver does not adopt the topology change.
Suggested fix
Poll under a bounded timeout until queries observe every expected host after node addition. Apply the same bounded eventual check after decommissioning so the test verifies that removed hosts disappear. Keep the topology logs as diagnostics rather than treating them as the readiness barrier.
Problem
ControlConnectionTests.Integration_Cassandra_TopologyChangeintermittently fails after bootstrapping node 2. The test expects queries to use two hosts, but sometimes observes only node 1:Recent failure:
https://github.com/scylladb/cpp-rs-driver/actions/runs/36589927177/job/109483566286
The identical assertion failed on
masterat commit6a4b16284469cbf001c8735e70b4c628cc2be496, before PR #509:https://github.com/scylladb/cpp-rs-driver/actions/runs/34943697382/job/104300208405
Cause
The test waits for the
Node added to clusterlog and then immediately makes onlyexpected_nodes.size() + 2requests:cpp-rs-driver/tests/src/integration/tests/test_control_connection.cpp
Lines 40 to 63 in c400fbc
cpp-rs-driver/tests/src/integration/tests/test_control_connection.cpp
Lines 318 to 346 in c400fbc
The Rust driver emits the node-added log while constructing refreshed cluster state, before it finishes initializing all pools and publishes that state for query routing. The log therefore confirms discovery, not readiness for requests. Immediate queries can still use the old one-node state.
Expected behavior
The integration test should tolerate asynchronous topology publication while still failing deterministically when the driver does not adopt the topology change.
Suggested fix
Poll under a bounded timeout until queries observe every expected host after node addition. Apply the same bounded eventual check after decommissioning so the test verifies that removed hosts disappear. Keep the topology logs as diagnostics rather than treating them as the readiness barrier.