ML-QuantSubscribe

arXivCrypto & DeFi

ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

Introduces a large-scale benchmark on 1.35 billion Ethereum transactions to evaluate DeFi representations under temporal, protocol, and contract drift, revealing model degradation on unseen pools.

Featured in No. 132 on 25 Sep 2026 · 3 days after release · 0 citations today

ETH-TraceBench benchmark construction and evaluation pipeline
Figure 1: ETH-TraceBench benchmark construction and evaluation pipeline. The raw Ethereum universe is converted into transaction-level event streams from receipt logs; DEX and liquidation labels are derived from curated DeFi tables; chronological partitions are fixed before model evaluation; and te…
Released
22 Sep 2026
First featured
No. 132 · 25 Sep 2026
Citations (Semantic Scholar)
0
Influential citations
0
Published in
Not yet, as far as Semantic Scholar knows
Fanfare
3 of 5
Identifier
arXiv:2609.23659
Authors
Kemal Kirtac and Carsten Maple

Abstract

From arXiv (CC0).

Ethereum decentralized finance (DeFi) provides a public, time-stamped record of transaction-level event streams, but the same public symbols can create strong machine-learning shortcuts. We introduce ETH-TraceBench, a benchmark for evaluating Ethereum DeFi representations under temporal, protocol, pool/infrastructure, and symbolic shift. The raw event universe covers January 2021-December 2025 and contains 1.35 billion transactions with logs and 5.01 billion raw log rows. Model evaluation uses a fixed 911,267-instance supervised sample, training on 2021-2024, selecting models on 2025H1, and testing on 2025H2. Simple models perform strongly on the aggregate temporal test: TraceStats-GB reaches 0.953 macro-F1 and TopicEmitterHashMLP 0.959 on the canonical DEX test set. Performance drops sharply under protocol novelty, with macro-F1 of 0.794, 0.743, and 0.766 for TraceStats-GB, TopicEmitterTrace-SGD, and TopicEmitterHashMLP, while strict unseen-pool scores remain 0.927, 0.897, and 0.935. Uniswap v4 and Ekubo v1, both absent from supervised training, are materially harder than the full test. Jointly masking emitter and topic identity reduces DEX macro-F1 to 0.916 and liquidation macro-F1 to 0.774 for TopicEmitterTrace-SGD. A standard Transformer over log-index-ordered events provides no consistent advantage over a deterministic shuffle of the same events, indicating that high aggregate scores can arise without sophisticated chronological modeling. A natural-prevalence audit estimates 2025H2 DEX prevalence among logged Ethereum transactions at about 22.5%, and a deterministic 400-transaction audit finds complete agreement with task label sources and independently re-queried raw-log counts. ETH-TraceBench therefore treats difficult transfer and controlled-input conditions, rather than a single aggregate score, as the main evaluation target.

Citations and venue from Semantic Scholar (ODC-BY), refreshed weekly. Summary: Quant Letter (CC BY 4.0).

    Type to search. Try rough volatility, LLM agents or FinGPT.

    ↑↓ move↵ openesc closeFull search page