Fixing OpenClaw Agent Crashes After 24 Hours
Fixing OpenClaw Agent Crashes After 24 Hours

If you're reading this, your OpenClaw agent probably died overnight and you woke up to either a cryptic error log or, worse, nothing at all. Just silence. The agent that was happily crunching through your task queue at midnight is now a dead process and you have no idea why.
I've been there. Multiple times. And after months of running OpenClaw agents in production — some processing thousands of tasks per day — I can tell you that the 24-hour crash is one of the most common and most fixable problems you'll encounter. It's almost a rite of passage.
The frustrating part isn't that it crashes. Software crashes. The frustrating part is that it works perfectly for hours before it does. You test it, you run it for an hour, everything's green, you go to bed confident, and then somewhere around hour 18 to 26, it falls over. Every single time.
Let me walk you through what's actually happening and how to fix it for good.
Why Your Agent Crashes After 24 Hours
There's almost never a single cause. It's usually a combination of three things compounding on each other over time, and they all hit critical mass around the same window:
1. Memory leaks from accumulated context 2. Stale or exhausted database/API connections 3. Unhandled rate limit or timeout errors that snowball
Let's break each one down.
The Memory Problem
This is the big one. If you're running an OpenClaw agent with default settings, every interaction, every tool call, every LLM response is being held in memory as part of the agent's context history. That's by design — your agent needs conversational memory to function well. But over 24 hours, that history becomes enormous.
Here's what typically happens: your agent starts at around 200-400MB of RAM. Reasonable. Over the next several hours, as it processes tasks, memory creeps up. By hour 12, you're at 2GB. By hour 20, you're at 8GB. Somewhere around hour 22-26, depending on your machine, you hit the ceiling and the OS kills the process. Or the garbage collector starts thrashing so hard that your agent can't respond to health checks, and your orchestrator kills it.
The default MemoryManager in OpenClaw doesn't aggressively prune because it's optimized for accuracy, not longevity. For short-running agents, that's fine. For anything meant to run continuously, you need to configure it explicitly.
Here's the configuration that fixed it for me:
from openclaw import Agent, MemoryManager
agent = Agent(
memory=MemoryManager(
max_context_tokens=8000,
compression='semantic',
eviction_policy='sliding_window',
vector_store_pool_size=5,
gc_interval=300 # Force garbage collection every 5 minutes
)
)
The two critical settings here are compression='semantic' and eviction_policy='sliding_window'.
Semantic compression means that instead of keeping every raw message in history, OpenClaw periodically summarizes older interactions into compressed representations. Your agent retains the meaning of what happened without holding every token. This alone can cut memory usage by 60-70% over long runs.
Sliding window eviction means the agent only keeps the N most recent interactions in full fidelity. Older interactions get compressed or dropped based on a relevance score. Without this, your context just grows unbounded.
You can also add explicit memory monitoring to catch problems before they crash your process:
@agent.task
async def process_with_memory_guard(item):
result = await agent.process(item)
if agent.memory.usage_percent > 80:
await agent.memory.compress()
logger.warning(f"Memory compressed at {agent.memory.usage_percent}%")
if agent.memory.usage_percent > 95:
await agent.memory.flush_non_essential()
logger.critical("Emergency memory flush triggered")
return result
This isn't elegant, but it works. I've had agents run for weeks straight with this pattern.
The Connection Pooling Problem
This one is sneakier. If your agent uses any external resources — vector stores like ChromaDB or Pinecone, databases, external APIs — each interaction potentially opens a new connection. Over 24 hours, those connections pile up.
Most connection libraries have default timeouts and pool sizes, but OpenClaw agents can easily blow through them because of how frequently they make calls. You open connections, some complete, some hang, some time out but don't close cleanly. Eventually you hit your pool limit or your file descriptor limit and everything stops.
The fix is explicit connection pooling and cleanup:
from openclaw import Agent, RobustExecutor
agent = Agent(
executor=RobustExecutor(
auto_retry=True,
max_retries=3,
exponential_backoff=True,
connection_pool_size=10,
connection_max_lifetime=3600, # Recycle connections every hour
connection_health_check=True
)
)
The connection_max_lifetime setting is the one most people miss. It forces connections to be recycled after a set period, even if they appear healthy. This prevents the slow accumulation of zombie connections that look fine but are actually degraded.
If you're using a vector store, set its pool size explicitly too:
agent = Agent(
memory=MemoryManager(
vector_store_pool_size=5, # Don't let this grow unbounded
vector_store_recycle_interval=1800 # Recycle every 30 minutes
)
)
I learned this the hard way. Had an agent that would crash at almost exactly the 22-hour mark, every time. Memory looked fine. CPU looked fine. Turned out it was exhausting ChromaDB connections. The error message was a generic ConnectionError with no additional context, which is why it took me three days to figure out.
The Cascade Failure Problem
The third piece of the puzzle is error handling — specifically, what happens when your agent hits a rate limit or timeout from an LLM provider.
Here's the scenario: your agent is humming along, makes a call that gets rate-limited, retries immediately, gets rate-limited again, retries again, now there's a backlog of tasks, each one making its own retry attempts, and suddenly your agent is making 10x the normal number of API calls, all of which are failing. CPU spikes. Memory spikes from all the queued responses. If you have a budget limit set, you blow through it. If you don't, you blow through your wallet.
OpenClaw has a circuit breaker pattern built in, but you have to enable it:
from openclaw import Agent, RobustExecutor
agent = Agent(
executor=RobustExecutor(
auto_retry=True,
max_retries=3,
exponential_backoff=True,
circuit_breaker_threshold=5, # Open circuit after 5 failures
circuit_breaker_timeout=60, # Wait 60 seconds before retrying
fallback_models=['gpt-4', 'gpt-3.5-turbo']
)
)
The circuit breaker stops your agent from hammering a failing endpoint. After a configurable number of consecutive failures, it "opens the circuit" — stops making calls entirely for a cooldown period. When the cooldown expires, it tries one test request. If that succeeds, normal operation resumes. If not, the circuit stays open for another cooldown.
Combined with fallback models, this means your agent can survive provider outages gracefully. If GPT-4 is rate-limited, it falls back to GPT-3.5-turbo. If everything is down, it pauses instead of burning through retries.
The Checkpointing Safety Net
Even with all of the above, things can still go wrong. Hardware fails. Providers have outages that last longer than your circuit breaker timeout. Your Kubernetes node gets rescheduled.
This is why checkpointing is non-negotiable for any agent running longer than an hour:
from openclaw import Agent, RobustExecutor
agent = Agent(
executor=RobustExecutor(
checkpoint_interval=50, # Save state every 50 operations
checkpoint_storage='disk', # Or 'redis', 's3'
auto_resume=True
)
)
@agent.task
async def process_documents(docs):
for doc in docs:
result = await agent.analyze(doc)
yield result # State auto-saved after each yield
With auto_resume=True, if your agent crashes and restarts, it picks up from the last checkpoint instead of starting over. This changes the crash from a catastrophic data loss event to a minor inconvenience — you lose at most 50 operations of progress.
I've had agents crash and resume so seamlessly that I didn't even notice until I checked the logs the next morning. That's the goal. Not preventing all crashes — that's impossible — but making crashes survivable.
Adding Observability So You Can See It Coming
Once you've fixed the immediate crash causes, you want visibility into your agent's health so you can catch problems before they become crashes:
from openclaw import Agent, Tracer
agent = Agent(
tracing=Tracer(
log_level='detailed',
capture_llm_calls=True,
capture_tool_calls=True,
cost_tracking=True,
metrics_export='prometheus' # Or 'datadog', 'cloudwatch'
)
)
The Tracer gives you structured logs with actual context — not walls of text, but JSON logs that tell you exactly which task was being processed, which tool was called, how many tokens were used, and what the cost was.
You can also set up budget controls so a runaway agent can't drain your API credits:
from openclaw import BudgetManager
agent = Agent(
budget=BudgetManager(
max_cost_per_task=5.00,
max_daily_cost=100.00,
alert_threshold=0.8
)
)
When the agent hits 80% of its daily budget, it fires an alert. When it hits 100%, it stops. No more $2,000 weekend surprise bills.
The Full Production Configuration
Here's the complete configuration I use for any long-running OpenClaw agent. This has kept agents stable for weeks at a time:
from openclaw import Agent, MemoryManager, RobustExecutor, Tracer, BudgetManager, ContextManager
agent = Agent(
memory=MemoryManager(
max_context_tokens=8000,
compression='semantic',
eviction_policy='sliding_window',
vector_store_pool_size=5,
gc_interval=300
),
executor=RobustExecutor(
auto_retry=True,
max_retries=3,
exponential_backoff=True,
checkpoint_interval=50,
checkpoint_storage='redis',
auto_resume=True,
circuit_breaker_threshold=5,
circuit_breaker_timeout=60,
fallback_models=['gpt-4', 'gpt-3.5-turbo'],
connection_pool_size=10,
connection_max_lifetime=3600
),
context=ContextManager(
strategy='priority_based',
max_tokens=8000,
preserve=['system_prompt', 'last_5_messages'],
summarize_old=True
),
tracing=Tracer(
log_level='detailed',
capture_llm_calls=True,
cost_tracking=True
),
budget=BudgetManager(
max_cost_per_task=5.00,
max_daily_cost=100.00,
alert_threshold=0.8
)
)
It looks like a lot of configuration, and honestly, it is. This is the part where I'll be straightforward with you: setting all of this up manually, testing each parameter, and tuning it for your specific use case takes time. I spent weeks getting this right through trial and error.
If you don't want to go through that entire process yourself, Felix's OpenClaw Starter Pack on Claw Mart is genuinely worth the $29. It includes pre-configured skills with production-hardened defaults for exactly these problems — memory management, connection pooling, checkpointing, circuit breaking, all of it. You get configurations that someone has already battle-tested in long-running scenarios, plus a set of pre-built skills that handle the most common agent patterns. I wish it had existed when I was figuring all of this out manually. It would have saved me a couple of very frustrating weeks.
Quick Diagnostic Checklist
When your agent crashes after extended runtime, run through this list:
Check memory first. Look at RSS memory over time. If it's a straight line up, you have a leak. Enable semantic compression and sliding window eviction.
Check connections second. Look at open file descriptors or connection counts. If they're climbing, enable connection pooling and lifecycle management.
Check error rates third. If you see bursts of errors followed by the crash, you're hitting cascade failures. Enable the circuit breaker.
Check disk/storage. If you're checkpointing or logging, make sure you're not filling up the disk. Set log rotation and checkpoint retention policies.
Check token usage. If your cost per task is climbing over time, your context is growing and the agent is sending more tokens per request. That's the context window filling up. Enable ContextManager with priority-based eviction.
Nine times out of ten, the 24-hour crash is a memory issue. The other one time, it's connections. Start there.
What to Do Next
If your agent is crashing right now, do this in order:
-
Add
MemoryManagerwithcompression='semantic'andeviction_policy='sliding_window'. Restart the agent. See if it survives longer. -
If it still crashes but later, add
RobustExecutorwith connection pooling and circuit breaking. -
Add checkpointing so crashes don't lose progress.
-
Add
TracerandBudgetManagerfor visibility and cost protection. -
Once stable, tune the parameters. Maybe you can increase
max_context_tokensfor better accuracy, or decreasecheckpoint_intervalfor tighter recovery.
Long-running OpenClaw agents are incredibly powerful once they're stable. Getting them stable is a one-time investment in configuration that pays off every single night you go to sleep without worrying about what you'll find in the morning. Set it up right once, and the crashes stop.
Recommended for this post