ClawMart AI
← Back to Blog
September 9, 20268 min readClaw Mart Team

Fixing OpenClaw Agent Crashes After 24 Hours

Fixing OpenClaw Agent Crashes After 24 Hours

Fixing OpenClaw Agent Crashes After 24 Hours

If you're reading this, your OpenClaw agent probably died overnight and you woke up to either a cryptic error log or, worse, nothing at all. Just silence. The agent that was happily crunching through your task queue at midnight is now a dead process and you have no idea why.

I've been there. Multiple times. And after months of running OpenClaw agents in production — some processing thousands of tasks per day — I can tell you that the 24-hour crash is one of the most common and most fixable problems you'll encounter. It's almost a rite of passage.

The frustrating part isn't that it crashes. Software crashes. The frustrating part is that it works perfectly for hours before it does. You test it, you run it for an hour, everything's green, you go to bed confident, and then somewhere around hour 18 to 26, it falls over. Every single time.

Let me walk you through what's actually happening and how to fix it for good.

Why Your Agent Crashes After 24 Hours

There's almost never a single cause. It's usually a combination of three things compounding on each other over time, and they all hit critical mass around the same window:

1. Memory leaks from accumulated context 2. Stale or exhausted database/API connections 3. Unhandled rate limit or timeout errors that snowball

Let's break each one down.

The Memory Problem

This is the big one. If you're running an OpenClaw agent with default settings, every interaction, every tool call, every LLM response is being held in memory as part of the agent's context history. That's by design — your agent needs conversational memory to function well. But over 24 hours, that history becomes enormous.

Here's what typically happens: your agent starts at around 200-400MB of RAM. Reasonable. Over the next several hours, as it processes tasks, memory creeps up. By hour 12, you're at 2GB. By hour 20, you're at 8GB. Somewhere around hour 22-26, depending on your machine, you hit the ceiling and the OS kills the process. Or the garbage collector starts thrashing so hard that your agent can't respond to health checks, and your orchestrator kills it.

The default MemoryManager in OpenClaw doesn't aggressively prune because it's optimized for accuracy, not longevity. For short-running agents, that's fine. For anything meant to run continuously, you need to configure it explicitly.

Here's the configuration that fixed it for me:

from openclaw import Agent, MemoryManager

agent = Agent(
    memory=MemoryManager(
        max_context_tokens=8000,
        compression='semantic',
        eviction_policy='sliding_window',
        vector_store_pool_size=5,
        gc_interval=300  # Force garbage collection every 5 minutes
    )
)

The two critical settings here are compression='semantic' and eviction_policy='sliding_window'.

Semantic compression means that instead of keeping every raw message in history, OpenClaw periodically summarizes older interactions into compressed representations. Your agent retains the meaning of what happened without holding every token. This alone can cut memory usage by 60-70% over long runs.

Sliding window eviction means the agent only keeps the N most recent interactions in full fidelity. Older interactions get compressed or dropped based on a relevance score. Without this, your context just grows unbounded.

You can also add explicit memory monitoring to catch problems before they crash your process:

@agent.task
async def process_with_memory_guard(item):
    result = await agent.process(item)
    
    if agent.memory.usage_percent > 80:
        await agent.memory.compress()
        logger.warning(f"Memory compressed at {agent.memory.usage_percent}%")
    
    if agent.memory.usage_percent > 95:
        await agent.memory.flush_non_essential()
        logger.critical("Emergency memory flush triggered")
    
    return result

This isn't elegant, but it works. I've had agents run for weeks straight with this pattern.

The Connection Pooling Problem

This one is sneakier. If your agent uses any external resources — vector stores like ChromaDB or Pinecone, databases, external APIs — each interaction potentially opens a new connection. Over 24 hours, those connections pile up.

Most connection libraries have default timeouts and pool sizes, but OpenClaw agents can easily blow through them because of how frequently they make calls. You open connections, some complete, some hang, some time out but don't close cleanly. Eventually you hit your pool limit or your file descriptor limit and everything stops.

The fix is explicit connection pooling and cleanup:

from openclaw import Agent, RobustExecutor

agent = Agent(
    executor=RobustExecutor(
        auto_retry=True,
        max_retries=3,
        exponential_backoff=True,
        connection_pool_size=10,
        connection_max_lifetime=3600,  # Recycle connections every hour
        connection_health_check=True
    )
)

The connection_max_lifetime setting is the one most people miss. It forces connections to be recycled after a set period, even if they appear healthy. This prevents the slow accumulation of zombie connections that look fine but are actually degraded.

If you're using a vector store, set its pool size explicitly too:

agent = Agent(
    memory=MemoryManager(
        vector_store_pool_size=5,  # Don't let this grow unbounded
        vector_store_recycle_interval=1800  # Recycle every 30 minutes
    )
)

I learned this the hard way. Had an agent that would crash at almost exactly the 22-hour mark, every time. Memory looked fine. CPU looked fine. Turned out it was exhausting ChromaDB connections. The error message was a generic ConnectionError with no additional context, which is why it took me three days to figure out.

The Cascade Failure Problem

The third piece of the puzzle is error handling — specifically, what happens when your agent hits a rate limit or timeout from an LLM provider.

Here's the scenario: your agent is humming along, makes a call that gets rate-limited, retries immediately, gets rate-limited again, retries again, now there's a backlog of tasks, each one making its own retry attempts, and suddenly your agent is making 10x the normal number of API calls, all of which are failing. CPU spikes. Memory spikes from all the queued responses. If you have a budget limit set, you blow through it. If you don't, you blow through your wallet.

OpenClaw has a circuit breaker pattern built in, but you have to enable it:

from openclaw import Agent, RobustExecutor

agent = Agent(
    executor=RobustExecutor(
        auto_retry=True,
        max_retries=3,
        exponential_backoff=True,
        circuit_breaker_threshold=5,  # Open circuit after 5 failures
        circuit_breaker_timeout=60,   # Wait 60 seconds before retrying
        fallback_models=['gpt-4', 'gpt-3.5-turbo']
    )
)

The circuit breaker stops your agent from hammering a failing endpoint. After a configurable number of consecutive failures, it "opens the circuit" — stops making calls entirely for a cooldown period. When the cooldown expires, it tries one test request. If that succeeds, normal operation resumes. If not, the circuit stays open for another cooldown.

Combined with fallback models, this means your agent can survive provider outages gracefully. If GPT-4 is rate-limited, it falls back to GPT-3.5-turbo. If everything is down, it pauses instead of burning through retries.

The Checkpointing Safety Net

Even with all of the above, things can still go wrong. Hardware fails. Providers have outages that last longer than your circuit breaker timeout. Your Kubernetes node gets rescheduled.

This is why checkpointing is non-negotiable for any agent running longer than an hour:

from openclaw import Agent, RobustExecutor

agent = Agent(
    executor=RobustExecutor(
        checkpoint_interval=50,  # Save state every 50 operations
        checkpoint_storage='disk',  # Or 'redis', 's3'
        auto_resume=True
    )
)

@agent.task
async def process_documents(docs):
    for doc in docs:
        result = await agent.analyze(doc)
        yield result  # State auto-saved after each yield

With auto_resume=True, if your agent crashes and restarts, it picks up from the last checkpoint instead of starting over. This changes the crash from a catastrophic data loss event to a minor inconvenience — you lose at most 50 operations of progress.

I've had agents crash and resume so seamlessly that I didn't even notice until I checked the logs the next morning. That's the goal. Not preventing all crashes — that's impossible — but making crashes survivable.

Adding Observability So You Can See It Coming

Once you've fixed the immediate crash causes, you want visibility into your agent's health so you can catch problems before they become crashes:

from openclaw import Agent, Tracer

agent = Agent(
    tracing=Tracer(
        log_level='detailed',
        capture_llm_calls=True,
        capture_tool_calls=True,
        cost_tracking=True,
        metrics_export='prometheus'  # Or 'datadog', 'cloudwatch'
    )
)

The Tracer gives you structured logs with actual context — not walls of text, but JSON logs that tell you exactly which task was being processed, which tool was called, how many tokens were used, and what the cost was.

You can also set up budget controls so a runaway agent can't drain your API credits:

from openclaw import BudgetManager

agent = Agent(
    budget=BudgetManager(
        max_cost_per_task=5.00,
        max_daily_cost=100.00,
        alert_threshold=0.8
    )
)

When the agent hits 80% of its daily budget, it fires an alert. When it hits 100%, it stops. No more $2,000 weekend surprise bills.

The Full Production Configuration

Here's the complete configuration I use for any long-running OpenClaw agent. This has kept agents stable for weeks at a time:

from openclaw import Agent, MemoryManager, RobustExecutor, Tracer, BudgetManager, ContextManager

agent = Agent(
    memory=MemoryManager(
        max_context_tokens=8000,
        compression='semantic',
        eviction_policy='sliding_window',
        vector_store_pool_size=5,
        gc_interval=300
    ),
    executor=RobustExecutor(
        auto_retry=True,
        max_retries=3,
        exponential_backoff=True,
        checkpoint_interval=50,
        checkpoint_storage='redis',
        auto_resume=True,
        circuit_breaker_threshold=5,
        circuit_breaker_timeout=60,
        fallback_models=['gpt-4', 'gpt-3.5-turbo'],
        connection_pool_size=10,
        connection_max_lifetime=3600
    ),
    context=ContextManager(
        strategy='priority_based',
        max_tokens=8000,
        preserve=['system_prompt', 'last_5_messages'],
        summarize_old=True
    ),
    tracing=Tracer(
        log_level='detailed',
        capture_llm_calls=True,
        cost_tracking=True
    ),
    budget=BudgetManager(
        max_cost_per_task=5.00,
        max_daily_cost=100.00,
        alert_threshold=0.8
    )
)

It looks like a lot of configuration, and honestly, it is. This is the part where I'll be straightforward with you: setting all of this up manually, testing each parameter, and tuning it for your specific use case takes time. I spent weeks getting this right through trial and error.

If you don't want to go through that entire process yourself, Felix's OpenClaw Starter Pack on Claw Mart is genuinely worth the $29. It includes pre-configured skills with production-hardened defaults for exactly these problems — memory management, connection pooling, checkpointing, circuit breaking, all of it. You get configurations that someone has already battle-tested in long-running scenarios, plus a set of pre-built skills that handle the most common agent patterns. I wish it had existed when I was figuring all of this out manually. It would have saved me a couple of very frustrating weeks.

Quick Diagnostic Checklist

When your agent crashes after extended runtime, run through this list:

Check memory first. Look at RSS memory over time. If it's a straight line up, you have a leak. Enable semantic compression and sliding window eviction.

Check connections second. Look at open file descriptors or connection counts. If they're climbing, enable connection pooling and lifecycle management.

Check error rates third. If you see bursts of errors followed by the crash, you're hitting cascade failures. Enable the circuit breaker.

Check disk/storage. If you're checkpointing or logging, make sure you're not filling up the disk. Set log rotation and checkpoint retention policies.

Check token usage. If your cost per task is climbing over time, your context is growing and the agent is sending more tokens per request. That's the context window filling up. Enable ContextManager with priority-based eviction.

Nine times out of ten, the 24-hour crash is a memory issue. The other one time, it's connections. Start there.

What to Do Next

If your agent is crashing right now, do this in order:

  1. Add MemoryManager with compression='semantic' and eviction_policy='sliding_window'. Restart the agent. See if it survives longer.

  2. If it still crashes but later, add RobustExecutor with connection pooling and circuit breaking.

  3. Add checkpointing so crashes don't lose progress.

  4. Add Tracer and BudgetManager for visibility and cost protection.

  5. Once stable, tune the parameters. Maybe you can increase max_context_tokens for better accuracy, or decrease checkpoint_interval for tighter recovery.

Long-running OpenClaw agents are incredibly powerful once they're stable. Getting them stable is a one-time investment in configuration that pays off every single night you go to sleep without worrying about what you'll find in the morning. Set it up right once, and the crashes stop.

Recommended for this post

Run AI coding agents in persistent tmux sessions that survive crashes, retry on failure, and notify on completion.

OpenClawEngineering
CI
Clawgear IO
$9Buy

Claw Mart Daily

Get one AI agent tip every morning

Free daily tips to make your OpenClaw agent smarter. No spam, unsubscribe anytime.

More From the Blog