AEM Workflow Debugging — 6.5 LTS / AMS
Production-grade debugging for AEM Granite Workflow engine, launcher, Inbox, Sling Jobs, thread pools, and purge on AEM 6.5 LTS and Adobe Managed Services (AMS).
Variant Scope
- This skill is 6.5-lts-only (includes AMS).
- Full JMX access via Felix Console or JMX client.
- Config changes via Felix Console or OSGi config in repository.
When to use this skill
- Workflow stuck, not progressing, failed, not starting, task not in Inbox, purge/repository bloat, permissions, queue backlog, thread pool exhaustion, auto-advancement not working.
- User provides thread dumps, configuration status ZIPs, Sling Job console output, or error.log excerpts.
- Environment: AEM 6.5 LTS / AMS (JMX available).
Step 1: Map symptom to first action
Step 2: Decision tree (workflow stuck)
- No current work item? → Stale. JMX:
countStaleWorkflows→restartStaleWorkflows(dryRun=true). - Participant step → Assignee exists? Inbox visible? Payload accessible? Dynamic participant resolver returning correct user?
- Process step → Search error.log for instance ID. Check:
process.labelregistered, payload path exists, bundle active, no exception inexecute(). - OR/AND Split → Condition evaluates correctly? Routes exist? No dead-end branches? Model synced?
Step 3: Thread dump & thread pool analysis
Thread dumps on 6.5 / AMS are obtained via jstack or by requesting from AMS support. Configuration status ZIPs from Felix Console → Status → Configuration Status.
3a. Sling default thread pool (critical path)
The Sling Scheduler ApacheSlingdefault uses ThreadPool: default. This pool runs:
- Oak observation events
- All Quartz-scheduled jobs — including the workflow timeout-detection scheduler that emits
com/adobe/granite/workflow/timeout/jobevents to the Sling Job system (the job itself then runs on the Granite Workflow Queue, see Step 3c)
Check the Sling Thread Pools status page (/system/console/status-slingthreadpools):
If active count = max pool size AND block policy = ABORT:
- New scheduled tasks (including workflow timeout/auto-advance jobs) are silently rejected
- This is the #1 cause of auto-advancement failure
Check the Threads status page (/system/console/status-Threads) or the jstack thread dump (/system/console/status-jstack-threaddump):
- Search for
sling-default-threads - If all threads show same stack (e.g. stuck on HTTP call, database, or external service), that's the blocking culprit
- Note
elapsedtime — threads stuck for hours indicate a hung external call without timeout
3b. Sling Job thread pool
Check Apache Sling Job Thread Pool in the Sling Thread Pools status page:
- active count vs max pool size
- If saturated, Sling Jobs cannot execute (workflow jobs stall)
3c. Granite Workflow Queue
Check the Sling Jobs page (/system/console/slingevent):
Check topic statistics for workflow model:
- Topic:
com/adobe/granite/workflow/job/var/workflow/models/<modelName> - High
Failed Jobs/ lowFinished Jobsratio → process step throwing exceptions
Check Granite Workflow Queue configuration:
- Type: Topic Round Robin
- Max Parallel: 0.5 OOTB on AEM 6.5 LTS (50% of available CPU cores). Increase for throughput on bursty workloads. Verify the running value at
/system/console/configMgr/org.apache.sling.event.jobs.QueueConfiguration~workflowbefore assuming. - Max Retries: 10
3d. Sling Scheduler
Check the Sling Scheduler status page (/system/console/status-slingscheduler):
- This page lists Quartz-style schedulers, not Sling Job topics. On OOTB AEM 6.5 LTS the workflow-related entry visible here is the periodic
WorkflowStatsMBeancollector (used by the Statistics MBean) — itsnextFireTimeshould be in the near future;nextFireTime: nullmeans the trigger was deregistered. - The
com/adobe/granite/workflow/timeout/jobtopic itself is a Sling Job, not a Quartz job — check it on the Sling Jobs page (/system/console/slingevent), not here. - Confirm
ApacheSlingdefaultusesThreadPool: default— that's how the periodic timeout-detection scheduler reaches the workflow engine.
Step 4: Error log patterns
Step 5: Configuration checklist
In Felix Console → OSGi → Configuration (/system/console/configMgr):
Step 6: Remediation quick reference
Step 7: Key JMX MBeans
All workflow maintenance and diagnostic operations live on a single MBean: com.adobe.granite.workflow:type=Maintenance. A separate MBean — com.adobe.granite.workflow:type=Statistics — exposes time-series workflow execution metrics for trend analysis.
Always use dryRun=true first before executing destructive purge or retry operations.
Step 8: Common root cause patterns (from real incidents)
Pattern A: Thread pool starvation → auto-advance failure
Symptom: Workflow auto-advancement stops; timeout jobs not firing; workflows stuck at participant step despite timeout configured.
Root cause chain:
- Custom scheduler (e.g.
AccessTokenScheduler) makes blocking HTTP call without timeout concurrent = trueallows overlapping executions on each cron trigger- Each stuck execution consumes a
defaultpool thread indefinitely - All pool threads consumed (OOTB on AEM 6.5 LTS that's only 5; environments hardened for throughput typically run with 20+) → pool saturated
- If block policy has been changed to
ABORT(OOTB isRUN), new Quartz triggers are rejected silently; onRUNthey instead pile up on the caller thread and back-pressure the dispatch - The workflow timeout-detection scheduler cannot dispatch new
com/adobe/granite/workflow/timeout/jobevents - Auto-advancement never happens
Diagnosis checklist:
- Sling Thread Pools page (
/system/console/status-slingthreadpools): Pooldefault→ active count = max pool size? - Sling Thread Pools page: Pool
default→ block policy = ABORT? - Threads page (
/system/console/status-Threads) or jstack: Allsling-default-*threads stuck on same stack? - Sling Jobs page (
/system/console/slingevent): Workflow job topic has high Failed Jobs? - Sling Scheduler page (
/system/console/status-slingscheduler): ThreadPool =defaultforApacheSlingdefault?
Fix: Restart instance (immediate); fix scheduler code (add HTTP timeout, set concurrent=false); change pool policy to RUN; increase pool size.
Pattern B: High workflow job failure rate
Symptom: numberOfFailedJobs >> numberOfFinishedJobs for a workflow topic.
Root cause: Process step exception, payload deleted, or process not registered.
Diagnosis: Search error.log for Error executing workflow step + model name. Check process.label in Felix Console → OSGi Components.
Pattern C: Stale workflows accumulating
Symptom: Workflows in RUNNING state but no work items; Inbox empty despite running instances.
Root cause: Cannot archive workitem during transition; JCR session crash during step completion.
Diagnosis: Search for Cannot archive workitem; JMX countStaleWorkflows; restartStaleWorkflows(dryRun=true).
References
- For runbook locations: see reference.md [blocked]


