Skip to main content

Performance Tuning

Optimize Codex for large libraries, high traffic, or resource-constrained environments.

Library Size Guidelines​

Library SizeRecommended SetupDatabaseWorkers
< 1,000 booksSingle instanceSQLite2-4
1,000 - 10,000 booksSingle instanceSQLite or PostgreSQL4
10,000 - 50,000 booksSingle instancePostgreSQL4-8
50,000+ booksMultiple instancesPostgreSQL8+ per instance

Database Optimization​

SQLite Performance​

SQLite works well for small to medium libraries with few concurrent users.

Enable WAL Mode​

Write-Ahead Logging significantly improves concurrent read/write performance:

database:
db_type: sqlite
sqlite:
path: ./data/codex.db
pragmas:
journal_mode: WAL
synchronous: NORMAL
tip

WAL mode allows readers while a write is in progress, dramatically improving responsiveness during library scans.

Connection Pool Behaviour Under Load​

Two settings keep a single SQLite instance from stalling when many requests arrive at once (for example, several open browser tabs):

  • batch_fan_out (default 4) bounds how many related-table queries a single list/detail request runs concurrently. Without this bound, one request could hold a connection per query and a few simultaneous requests would exhaust the pool, causing multi-second connection pool timeout waits.
  • background_max_connections (default 4) gives in-process task workers, the scheduler, and pollers a separate connection pool. A library scan or analysis burst therefore cannot drain the pool that serves interactive API requests. This pool is only created when workers run in the same process (the default single-container serve); web-only deployments keep a single pool.

If you still hit pool-timeout warnings under heavy concurrent browsing, raise max_connections (cheap under WAL — connections are just file handles).

note

These settings reduce connection contention, but writes still serialize on SQLite: a heavy scan slows other writes regardless of pool sizing. That is inherent to SQLite. For write-heavy or many-user workloads, use PostgreSQL.

SQLite Limitations​

  • Single writer: Only one process can write at a time (heavy scans slow all writes)
  • No horizontal scaling: Cannot run multiple Codex instances
  • Concurrent users: Best for 5-10 simultaneous users
  • File locking: Network filesystems (NFS, SMB) may cause issues

PostgreSQL Performance​

PostgreSQL is recommended for larger libraries and multiple concurrent users.

Connection Pooling​

Configure connection limits based on your workload:

database:
db_type: postgres
postgres:
host: localhost
port: 5432
username: codex
password: your-password
database_name: codex
max_connections: 25
min_connections: 2
connect_timeout: 30
idle_timeout: 600

Guidelines:

  • max_connections is a per-process ceiling. Divide the server's budget between the processes that share it: (server max_connections - superuser_reserved_connections) / peak process count, counting every web replica, worker, migration Job and backup CronJob, plus the surge pods a rolling update briefly adds.
  • min_connections: keep it small. Warm connections are held per process whether or not it is busy, so this multiplies by replica count.
  • Raise max_connections for throughput only once you have confirmed the total still fits. Overrunning it stops new pods from starting, since a process needs a connection before it can do anything at all.

PostgreSQL Tuning​

For large libraries, tune PostgreSQL itself:

# postgresql.conf

# Memory (adjust based on available RAM)
shared_buffers = 256MB # 25% of RAM, up to 1GB
effective_cache_size = 768MB # 75% of RAM
work_mem = 16MB # Per-operation memory

# Write Performance
wal_buffers = 16MB
checkpoint_completion_target = 0.9

# Query Planning
random_page_cost = 1.1 # For SSDs
effective_io_concurrency = 200 # For SSDs

Index Maintenance​

Periodically optimize indexes:

# Vacuum and analyze (run weekly)
psql -U codex codex -c "VACUUM ANALYZE"

# Reindex if queries slow down (run monthly)
psql -U codex codex -c "REINDEX DATABASE codex"

Worker Configuration​

Background workers handle scanning, thumbnail generation, and analysis tasks.

Worker Count​

task:
worker_count: 4

Guidelines:

  • Minimum: 2 workers (one for scanning, one for other tasks)
  • Recommended: CPU cores - 1 (leave headroom for web server)
  • Maximum: 2 × CPU cores (diminishing returns beyond this)
caution

More workers isn't always better. Too many workers can cause:

  • Memory exhaustion (each worker loads files into memory)
  • Disk I/O saturation
  • Database connection pool exhaustion

Scanner Concurrency​

Limit simultaneous library scans:

scanner:
max_concurrent_scans: 2

Scanning is I/O intensive. Running too many concurrent scans can:

  • Saturate disk bandwidth
  • Cause memory pressure from parallel file processing
  • Slow down the web interface

Memory Optimization​

Thumbnail Configuration​

Thumbnails are the largest memory consumers. Configure via the Settings API:

# Reduce thumbnail size for memory-constrained systems
curl -X PATCH http://localhost:8080/api/v1/admin/settings \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"thumbnail_max_dimension": 300,
"thumbnail_jpeg_quality": 75
}'
SettingDefaultLow MemoryHigh Quality
thumbnail_max_dimension400200-300500-600
thumbnail_jpeg_quality8570-7590-95

Preload Settings​

In large libraries, preloading can consume significant memory:

# codex.yaml - no preload settings, these are client-side

Preload settings are per-user (stored in browser). For memory-constrained servers, advise users to reduce preload pages in Reader Settings.

Scanning Performance​

Initial Scan Strategy​

For large libraries, consider:

  1. Scan in batches: Add libraries one at a time
  2. Use normal mode first: Deep scans are more resource-intensive
  3. Schedule during off-hours: Use cron scheduling
# Example: Library with cron schedule for nightly deep scans
# Configure via UI when creating library

File System Considerations​

Storage TypePerformanceRecommendation
Local SSDExcellentBest performance
Local HDDGoodAcceptable for small libraries
Network (NFS)VariableUse PostgreSQL, may need tuning
Network (SMB)PoorNot recommended for large libraries

:::warning Network Storage Network-mounted libraries may experience:

  • Slower scanning (metadata extraction over network)
  • File locking issues with SQLite
  • Timeout errors for large files

Use PostgreSQL and increase timeouts for network storage. :::

Duplicate Detection​

Hash-based duplicate detection is CPU and I/O intensive:

  • Enable for: Libraries with potential duplicates
  • Disable for: Clean, well-organized libraries
  • Schedule: Run duplicate scans during low-usage periods

Horizontal Scaling​

For high-traffic deployments, run multiple Codex instances behind a load balancer.

Requirements​

  • PostgreSQL: Required (SQLite doesn't support multiple writers)
  • Shared storage: All instances must access the same media files
  • Shared thumbnails: Mount thumbnail directory on shared storage or use object storage

Architecture​

┌─────────────────┐
│ Load Balancer │
│ (Nginx/HAProxy)│
└────────┬────────┘
│
┌─────────────────┼─────────────────┐
│ │ │
▼ ▼ ▼
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Codex 1 │ │ Codex 2 │ │ Codex 3 │
│ (workers) │ │ (workers) │ │ (workers) │
└──────┬──────┘ └──────┬──────┘ └──────┬──────┘
│ │ │
└─────────────────┼─────────────────┘
│
┌────────┴────────┐
│ PostgreSQL │
└─────────────────┘
│
┌────────┴────────┐
│ Shared Storage │
│ (Media + Thumbs)│
└─────────────────┘

Configuration for Scaling​

# Each instance
database:
db_type: postgres
postgres:
host: postgres-service
# ... connection settings

# Reduce workers per instance when scaling horizontally
task:
worker_count: 2

# Limit concurrent scans across cluster
scanner:
max_concurrent_scans: 1

Load Balancer Configuration​

# nginx.conf
upstream codex_cluster {
least_conn; # Route to least busy server
server codex1:8080;
server codex2:8080;
server codex3:8080;
}

server {
listen 80;
server_name library.example.com;

location / {
proxy_pass http://codex_cluster;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}

# SSE requires long timeouts
location /api/v1/sse {
proxy_pass http://codex_cluster;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_buffering off;
proxy_read_timeout 86400s;
}
}

Monitoring Performance​

Key Metrics​

Monitor these metrics for performance issues:

MetricHealthy RangeAction if High
Response time< 200msCheck database, add caching
Memory usage< 80%Reduce workers, thumbnail size
CPU usage< 70%Reduce workers, scan concurrency
Database connections< 80% of maxIncrease pool or reduce workers
Disk I/O wait< 20%Use faster storage, reduce concurrency

Health Endpoint​

# Basic health check
curl http://localhost:8080/health

# Detailed metrics (requires authentication)
curl http://localhost:8080/api/v1/metrics \
-H "Authorization: Bearer $TOKEN"

Slow Query Detection​

For PostgreSQL, enable slow query logging:

# postgresql.conf
log_min_duration_statement = 100 # Log queries > 100ms

Troubleshooting Performance​

Slow Library Scans​

  1. Check disk I/O: iostat -x 1
  2. Reduce concurrent scans
  3. Use normal mode instead of deep scan
  4. Check for network storage latency

High Memory Usage​

  1. Reduce worker count
  2. Lower thumbnail dimensions
  3. Check for memory leaks (restart periodically)
  4. Increase system swap as safety net

Slow Page Loads​

  1. Check database query times
  2. Verify thumbnail cache is working
  3. Check network latency to clients
  4. Enable response compression in reverse proxy

Database Connection Exhaustion​

Error: too many connections
  1. Reduce max_connections in Codex config
  2. Increase PostgreSQL max_connections
  3. Reduce worker count
  4. Check for connection leaks (long-running queries)