Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,9 +160,9 @@ Top-level layout follows the system's data flow. Each sequencer module correspon
- `l2_tx_feed/` — DB-backed ordered-tx feed.
- `sequencer/src/l1/` — L1 client surface.
- `reader.rs` — safe-input ingestion from InputBox into SQLite.
- `submitter/` — stateless batch submitter (`worker.rs` + `poster.rs`).
- `submitter/` — batch submitter (`worker.rs` + `poster.rs`); re-estimates fees every tick without carrying a fee floor from earlier attempts. The rationale, accepted liveness limits, and revisit criteria live in [`docs/l1-fee-policy.md`](docs/l1-fee-policy.md).
- `fee_oracle/` — setup-pinned L1 Uniswap V3 TWAP → `batch_policy.log_gas_price` (+ `log_gas_price_updated_at_ms`); fixed mode writes once at setup and has no worker.
- `eip1559.rs` — shared EIP-1559 fee estimation (poster + oracle).
- `eip1559.rs` — shared EIP-1559 fee estimation (poster, oracle, flusher).
- `provider.rs` — alloy provider construction.
- `partition.rs` — long-block-range retry helper.
- `sequencer/src/recovery/` — preemptive recovery startup procedure (`mod.rs`), runtime danger detector (`detector.rs`), and mempool flusher (`flusher.rs`).
Expand Down
74 changes: 74 additions & 0 deletions docs/l1-fee-policy.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# L1 fee policy

The batch poster uses a fresh market estimate on every tick, with no fee
floor carried from earlier attempts. This keeps retry frequency from
compounding the prices offered for pending batches. It accepts possible
prolonged submission stalls; prompt inclusion is not guaranteed.

## Submission

Each tick derives the unresolved batch suffix from persisted and observed
L1 state and attempts every payload, including those behind a refused
replacement. A remembered transaction hash never causes a payload to be
skipped. Gas is estimated without a nonce, at Latest, and padded before the
wallet nonce is attached to the send.

The [estimator](../sequencer/src/l1/eip1559.rs) uses ten historical blocks,
the median positive 20th-percentile reward as the priority cap, and
`fee cap = 2 × base fee + priority cap`. Geth's default replacement policy
requires both caps to increase strictly and meet its 10% bump threshold
(with integer rounding). A fresh estimate can improve one component without
clearing the other.

`already known` and `replacement transaction underpriced` can occur while
the original is still adequately priced. They do not by themselves establish
an inclusion problem. Recognized mempool conflicts and confirmation timeouts
produce `Waiting`; the worker sleeps its idle interval before re-estimating.
Other provider errors retain their existing error path.

## Accepted limits

A pending transaction can become unmineable when its fee cap falls below the
base fee, while fresh estimates still fail the replacement rule. It can also
remain mineable but uncompetitive. Full blocks do not guarantee a rising tip
estimate: base-fee changes are protocol rules, while tips reflect the fee
market and builder choices. Clearing within minutes is an operational
expectation to measure, not a bound this policy establishes.

The danger detector stops normal operation when the configured danger
threshold is reached. That bounds continued soft-confirmation issuance under
the detector's assumptions; it does not bound inclusion or recovery duration.
[Recovery](recovery/README.md#step-4-post-flush-state) requires all covered
wallet slots to resolve at safe depth before a cascade can proceed. The
flusher's fixed headroom can also fail to replace an unmineable original.
The sequencer remains offline until recovery succeeds or the operator acts.

The poster warns after a nonce has been unresolved for five minutes and at
most once per further interval. Accepted re-broadcasts do not reset its age.
The clock starts at this process's first attempt and resets on restart, so it
is a lower bound on unresolved age. A sequencer restart does not itself clear
the Ethereum node's pending transactions.

## Why this policy

The [#34 review history](https://github.com/cartesi/sequencer/pull/34) explored
asymmetric bumps, remembered suffix hashes, compounding fee floors, and a
market-relative ceiling. Those implementations introduced invalid fee pairs,
payload/hash-association bugs, funding pressure, and recovery incompatibility.
[#35](https://github.com/cartesi/sequencer/pull/35) retains the independent
gas-estimation and flusher fee-validity fixes while removing that escalation
state. These failures motivate the current choice; they do not prove that
every bounded replacement policy would be unsound.

Revisit with evidence of sustained batch delays, recovery frequency or
duration, and submission cost. Record the logs and affected L1 block range.
Any future escalation policy needs an explicit urgency trigger, spending
budget, and compatible recovery pricing. A rejected replacement alone is
insufficient motivation.

Fee-policy tests must distinguish [geth's two-component replacement
check](https://github.com/ethereum/go-ethereum/blob/master/core/txpool/legacypool/list.go)
from [Anvil's gas-price replacement
check](https://github.com/foundry-rs/foundry/blob/v1.4.3/crates/anvil/src/eth/pool/transactions.rs).
Anvil exercises the send and retry flow but does not establish geth's fee
acceptance behavior.
19 changes: 18 additions & 1 deletion docs/recovery/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -211,7 +211,24 @@ Every `w_nonce` slot from N to M-1 is now resolved:

There are no more mempool entries. All uncertainty is resolved.

**Flush safety does not depend on eviction.** A no-op may fail to evict a still-pending batch tx (e.g. our local node rejects the replacement under EIP-1559's ≥10% bump rule). That's fine: a rejected send surfaces as a hard `FlushError` and the process exits, the orchestrator respawn re-runs the flush, and *eventual* inclusion of either the original batch tx or the no-op resolves the slot — the unbounded retry lives in the respawn loop, not inside `flush_and_wait`. Safety holds regardless of which lands; eviction is only an operational efficiency concern.
**Flush safety does not depend on eviction; completion depends on L1 progress.**
A rejected no-op surfaces as a hard `FlushError` and the process exits. The
orchestrator respawn re-runs the flush. Inclusion of either the original batch
or a no-op can resolve the slot, but the sequencer remains offline until every
covered slot reaches safe depth. Neither retries nor the danger threshold
establish a recovery deadline.

No-ops use 3× the fresh fee estimate, followed by a symmetric replacement bump.
This headroom improves their chance of replacing an earlier transaction; it
does not guarantee replacement. Base fees and priority estimates can move in
opposite directions. For example, a poster tx sent at base 10 gwei with cap
22 gwei and tip 2 gwei cannot mine at base 30 gwei. If the current tip estimate
is 0.5 gwei, the no-op offers cap 199.65 gwei and tip 1.65 gwei (plus 1 wei on
each). Geth rejects that replacement because its tip misses the 2.2 gwei
threshold. Both the original and the no-op can therefore fail to make progress.
A previous flush no-op can also block another pass on a flat market. These are
accepted liveness limits of the current [fee policy](../l1-fee-policy.md), not
permission to cascade before the slots resolve.

### Step 5: Run recovery

Expand Down
18 changes: 17 additions & 1 deletion sequencer/src/l1/eip1559.rs
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
// (c) Cartesi and individual authors (see AUTHORS)
// SPDX-License-Identifier: Apache-2.0

//! Shared EIP-1559 fee estimate used by the poster and fee oracle.
//! Shared fee estimate and gas padding for the poster, oracle, and flusher.
//! The poster uses fresh estimates; the flusher applies fixed headroom.
//! Rationale and limitations: `docs/l1-fee-policy.md`.

use alloy::consensus::BlockHeader;
use alloy::providers::{DynProvider, Provider, utils};
Expand All @@ -15,6 +17,13 @@ pub struct Eip1559Fees {
pub max_fee_per_gas: u128,
}

/// Pad an `eth_estimateGas` result so a tight estimate cannot mine as an
/// out-of-gas revert (which would burn the wallet-nonce slot with no
/// `InputAdded` and desynchronize payload↔nonce assignment).
pub fn pad_gas_estimate(gas: u64) -> u64 {
gas.saturating_add(gas / 10)
}

/// Estimate fees with Alloy's default, MetaMask-style medium estimator.
///
/// We intentionally pin the policy constants here: 10 historical blocks, the
Expand Down Expand Up @@ -65,4 +74,11 @@ mod tests {
assert_eq!(estimate.max_priority_fee_per_gas, 4);
assert_eq!(estimate.max_fee_per_gas, 204);
}

#[test]
fn pad_gas_estimate_adds_ten_percent() {
assert_eq!(pad_gas_estimate(100_000), 110_000);
assert_eq!(pad_gas_estimate(0), 0);
assert_eq!(pad_gas_estimate(u64::MAX), u64::MAX);
}
}
43 changes: 32 additions & 11 deletions sequencer/src/l1/fee_oracle/worker.rs
Original file line number Diff line number Diff line change
Expand Up @@ -277,7 +277,8 @@ mod tests {
use super::*;
use crate::storage::test_helpers::temp_db;
use alloy_primitives::U256;
use std::sync::Mutex;
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::{Arc, Mutex};

const TEST_MAX_AGE_MS: u64 = 60 * 60 * 1000;

Expand All @@ -300,16 +301,15 @@ mod tests {
}

struct FailsAfterFirstGas {
calls: Mutex<usize>,
calls: Arc<AtomicUsize>,
ok: Eip1559Fees,
}

#[async_trait]
impl GasFeeSource for FailsAfterFirstGas {
async fn estimate_gas_fees(&self) -> Result<Eip1559Fees, String> {
let mut calls = self.calls.lock().expect("lock");
*calls += 1;
if *calls == 1 {
let n = self.calls.fetch_add(1, Ordering::SeqCst) + 1;
if n == 1 {
Ok(self.ok)
} else {
Err("rpc unavailable".into())
Expand Down Expand Up @@ -438,7 +438,7 @@ mod tests {
&db.path,
TEST_MAX_AGE_MS,
Box::new(FailsAfterFirstGas {
calls: Mutex::new(0),
calls: Arc::new(AtomicUsize::new(0)),
ok: sample_fees(),
}),
Box::new(StaticToken(sample_quote())),
Expand Down Expand Up @@ -532,23 +532,44 @@ mod tests {
initialize_db(&db.path);
let expected_log = expected_log_price();

let gas_calls = Arc::new(AtomicUsize::new(0));
let oracle = FeeOracle::new_with_sources(
db.path.clone(),
Duration::from_millis(40),
TEST_MAX_AGE_MS,
Box::new(FailsAfterFirstGas {
calls: Mutex::new(0),
calls: Arc::clone(&gas_calls),
ok: sample_fees(),
}),
Box::new(StaticToken(sample_quote())),
);
let shutdown = ShutdownSignal::default();
let mut handle = oracle.start(shutdown.clone());

tokio::select! {
biased;
result = &mut handle => panic!("fee oracle exited early: {result:?}"),
_ = tokio::time::sleep(Duration::from_millis(200)) => {}
// First refresh is a spawn_blocking SQLite write; a fixed sleep flakes
// when the blocking pool is busy. Wait until the price is persisted
// *and* a later tick has failed, so retain-on-transient is covered.
let deadline = tokio::time::Instant::now() + Duration::from_secs(2);
loop {
tokio::select! {
biased;
result = &mut handle => panic!("fee oracle exited early: {result:?}"),
_ = tokio::time::sleep(Duration::from_millis(10)) => {}
}
let price = Storage::open_read_only(&db.path)
.unwrap()
.log_gas_price()
.unwrap();
let calls = gas_calls.load(Ordering::SeqCst);
if price == expected_log && calls >= 2 {
break;
}
if tokio::time::Instant::now() >= deadline {
panic!(
"timed out waiting for retained fee-oracle price \
(expected {expected_log}, last read {price}, gas_calls={calls})"
);
}
}
let storage = Storage::open_read_only(&db.path).unwrap();
assert_eq!(storage.log_gas_price().unwrap(), expected_log);
Expand Down
3 changes: 2 additions & 1 deletion sequencer/src/l1/submitter/config.rs
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,8 @@ use std::time::Duration;
/// doesn't read it. The [`crate::recovery::DangerDetector`] worker owns that.
#[derive(Debug, Clone)]
pub struct BatchSubmitterConfig {
/// How often the submitter polls for new work when idle.
/// How often the submitter polls for new work when idle, and how long it
/// waits before re-estimating when the mempool already holds its txs.
pub idle_poll_interval_ms: u64,
}

Expand Down
5 changes: 4 additions & 1 deletion sequencer/src/l1/submitter/mod.rs
Original file line number Diff line number Diff line change
Expand Up @@ -15,5 +15,8 @@ mod poster;
mod worker;

pub use config::BatchSubmitterConfig;
pub use poster::{BatchPoster, BatchPosterConfig, BatchPosterError, EthereumBatchPoster, TxHash};
pub use poster::{
BatchPoster, BatchPosterConfig, BatchPosterError, EthereumBatchPoster, SubmitBatchesOutcome,
TxHash,
};
pub use worker::{BatchSubmitter, BatchSubmitterError, SubmitterExit};
Loading