Skip to main content

Optimizing system configuration

The number of possible setups is virtually unlimited, so this document cannot cover them all. Even similar builds and configurations can behave differently because of external factors. Your results may vary. Use these general guidelines as a starting point. Be cautious and make incremental changes. Test and observe each change before you move forward. Always focus on only one specific area at a time. Do not change the memory, storage, and CPU configurations all at once. Otherwise, it becomes nearly impossible to diagnose problems.

Memory management

These settings in /etc/sysctl.conf can optimize memory usage and disk I/O patterns:
To apply the changes, run sudo sysctl -p.

Network stack

These settings in /etc/sysctl.conf may improve network performance:

Storage configuration

For NVMe drives, optimize I/O scheduling: Storage optimization commands

Infrastructure monitoring

Monitoring is one of the most critical components of network infrastructure. This page covers monitoring, performance tuning, and alerting configuration for Cosmos SDK and Tendermint nodes.

Prometheus setup

First, install Prometheus:
This is an example Prometheus configuration:

Grafana integration

Install and configure Grafana:

Alert management

Install Alertmanager:

Log management

Loki setup

Use Loki for log aggregation:

Log rotation

Configure logrotate to manage log files:

Security configuration

Network security

Configure the UFW firewall:

Rate limiting

Validator-specific monitoring

Status query

Query the validator status through the SDK:
Query the status through the REST API:

Critical metrics

Monitor these validator-specific metrics:

Backup management

Host system monitoring

Resource usage tracking

Install and configure node_exporter:
Add node_exporter to the Prometheus configuration:

EVM RPC OpenTelemetry metrics

The EVM RPC layer emits OpenTelemetry metrics through the process-wide MeterProvider (for example, a Prometheus exporter). It emits these metrics in parallel with the legacy sei_* metrics, so you can migrate dashboards incrementally.

Available EVM RPC metrics

The evmrpc_request_latency_seconds histogram carries these labels:

Migrating from legacy EVM RPC metrics

These legacy sei_* metrics remain available today but are deprecated. They are scheduled for removal after dashboards migrate to the evmrpc_* OpenTelemetry metrics: Before the legacy metrics are removed, update your Prometheus and Grafana dashboards to use the evmrpc_* metrics. The latency unit changed from milliseconds (sei_rpc_request_latency_ms) to seconds (evmrpc_request_latency_seconds). Adjust any thresholds and panel formatting to match.

FlatKV OpenTelemetry metrics

The FlatKV state store emits OpenTelemetry metrics through the process-wide MeterProvider (for example, a Prometheus exporter). With these metrics, you can observe commit throughput, catchup progress, snapshotting, rollbacks, and snapshot imports.

Available FlatKV metrics

Labels

FlatKV metrics carry these labels where applicable:

Enabling Pebble internal metrics

A single FlatKV-level knob controls Pebble’s internal (per-DB) metrics: enable-pebble-metrics under [state-commit.flatkv] in app.toml (default true). The node honors this key when it is present, but the app.toml that seid init generates does not include it. To change the default, add the key manually. The value propagates to every data DB (account, code, storage, legacy, and metadata) during initialization. It overrides any per-DB EnableMetrics settings. Configure Pebble metrics through this knob, not through the individual per-DB settings.

LittDB OpenTelemetry metrics

LittDB emits its metrics through the process-wide OpenTelemetry MeterProvider instead of a private Prometheus client. When MetricsEnabled is set, LittDB configures a Prometheus exporter on the global provider and serves /metrics on MetricsPort (default 9101). The MetricsNamespace and MetricsRegistry config fields were removed, and all metric names use a fixed litt_ prefix.

Available LittDB metrics

Attributes

Migrating from legacy LittDB metrics

The OpenTelemetry migration changed the metric names, units, and shape. You must update your existing Prometheus and Grafana dashboards:
  • Latency metrics moved from millisecond summaries (for example, {namespace}_read_latency_ms) to second histograms (litt_read_latency_seconds). Change thresholds and panel formatting from milliseconds to seconds.
  • Counters and gauges gained a fixed litt_ prefix and explicit units. For example, bytes_read became litt_bytes_read, and the cache weight gauge became litt_chunk_cache_weight_bytes.
  • The per-cache series that previously had separate metric names (for example, chunk_read_cache_* and chunk_write_cache_*) are now the shared litt_chunk_cache_* metrics. The cache attribute distinguishes them.
  • The MetricsNamespace and MetricsRegistry config fields no longer exist. Metric names are fixed, and the global OTel provider always backs the metrics. Set the scrape port with MetricsPort.

Performance testing

For specific customizations or other metrics, ask the Sei technical communities on Telegram or Discord.