Using CloudWatch Metrics with AWS Batch
AWS Batch publishes job metrics to Amazon CloudWatch for monitoring state transitions and
durations of jobs as they move through their lifecycle. AWS Batch publishes these metrics
under the AWS/Batch namespace at no additional charge.
Metrics are emitted as jobs transition between states and as job attempts complete. State transition metrics track how many jobs entered a given state, and duration metrics track how long jobs or job attempts spent moving between states. Use these metrics to build CloudWatch dashboards, set CloudWatch alarms, and establish performance baselines for your workloads. For more information, see Using CloudWatch metrics and Using CloudWatch alarms in the Amazon CloudWatch User Guide.
The reporting behavior for Array jobs and Multi-node parallel jobs differs from single-node jobs. The Reporting behavior column in the following table notes these cases, and Metrics for array jobs and multi-node parallel jobs describes the array and MNP behavior in detail.
Note
AWS Batch emits these metrics for jobs submitted to a job queue. All metrics are
published with the JobQueueName dimension. You can use this dimension to
group and filter metric data by the job queue that a job was submitted to.
Important
These metrics are provided on a best-effort basis and are intended for monitoring and
observability rather than billing, auditing, or reconciliation. Because job state
transitions are event-driven, a metric can occasionally be missed. For example, if a job
begins running and then succeeds or fails almost immediately, the transition into the
RUNNING state might not be observed before the job reaches its terminal
state. In that case, AWS Batch does not emit the metrics associated with the
RUNNING transition for that job. Design any alarms or dashboards to
tolerate occasional gaps, and don't rely on exact counts derived from these
metrics.
AWS Batch CloudWatch metrics
AWS Batch sends the following job metrics to CloudWatch. State transition metrics use the
Count unit and duration metrics use the Milliseconds
unit.
The following metrics apply to AWS Batch jobs regardless of type.
Metric |
Description |
Units |
Reporting behavior |
|---|---|---|---|
|
The number of jobs that were submitted to the job queue. |
Count |
Emitted when a job is first submitted (enters the
|
|
The number of jobs that moved to the |
Count |
Emitted when a job moves to |
|
The number of jobs that moved to the |
Count |
Emitted when a job moves to |
|
The number of jobs that moved to the |
Count |
Emitted when a job moves to
|
|
The number of jobs that moved to the |
Count |
Emitted when a job moves to
|
|
The number of jobs that completed successfully. |
Count |
Emitted when a job moves to
|
|
The number of jobs that moved to the |
Count |
Emitted when a job enters the |
|
The number of jobs that were cancelled. |
Count |
Emitted when a job enters the |
|
The number of jobs that were terminated. |
Count |
Emitted when a job enters the |
|
The number of job attempts that were retried. |
Count |
Emitted when a job attempt fails and the job is returned to
|
|
The time a job took to go from |
Milliseconds |
Emitted when the job enters |
|
The time a job attempt spent in |
Milliseconds |
Emitted per attempt when the job enters |
|
The time a job attempt spent running, measured from when the attempt started to when it stopped. |
Milliseconds |
Emitted when a job attempt completes, measured from when the attempt started running to when it stopped. |
|
The total time a job attempt lasted, measured from when it
entered |
Milliseconds |
Emitted when an attempt ends. The attempt ends either when the
job reaches a terminal state ( |
The following metrics apply only to AWS Batch compute jobs.
Metric |
Description |
Units |
Reporting behavior |
|---|---|---|---|
|
The time a job took to go from |
Milliseconds |
Emitted on the first attempt, when the job first enters |
|
The time a job attempt spent in |
Milliseconds |
Emitted when the job enters |
The following metrics apply only to Service jobs in AWS Batch.
Metric |
Description |
Units |
Reporting behavior |
|---|---|---|---|
|
The number of jobs that moved to the |
Count |
Emitted for service jobs when the job first moves to
|
|
The time a service job took to go from |
Milliseconds |
Emitted for service jobs when the job first moves to
|
|
The time a job attempt spent in |
Milliseconds |
Emitted per attempt when a service job moves to
|
|
The time a service job attempt spent between first moving to
|
Milliseconds |
Emitted for service jobs when the job moves from
|
|
The number of jobs that were preempted. |
Count |
Emitted for quota management jobs each time the job is preempted. |
Metrics for array jobs and multi-node parallel jobs
AWS Batch represents Array jobs and Multi-node parallel jobs as a parent job that tracks one or more child jobs. So that a single logical job isn't counted twice, AWS Batch emits metrics from only one of these two records, depending on the job type.
- Array jobs
-
The array parent job does not emit metrics, except for
JobsSubmitted. Instead, each array child job emits these metrics independently as it moves through the job lifecycle, so the metric counts reflect the number of child jobs rather than the number of array jobs submitted.The
JobsSubmittedmetric is emitted once for the array parent job at submission, and its value is the size of the array (the number of child jobs). For example, submitting one array job with 1,000 children publishes a singleJobsSubmitteddata point with a value of1000.Duration metrics that measure time from submission — such as
JobSubmittedToRunnableDuration— are emitted per array child job and measure from when the array parent was submitted to when that child reached the state in question. - Multi-node parallel (MNP) jobs
-
MNP metrics are emitted from the MNP parent job, which represents the job as a whole. The individual MNP child (node) jobs do not emit these metrics. As a result, an MNP job is counted once regardless of how many nodes it uses.
Dimensions for AWS Batch CloudWatch metrics
AWS Batch job metrics in CloudWatch use a single dimension: JobQueueName. All
metric data is grouped and filtered by the name of the job queue that a job was
submitted to.
View AWS Batch CloudWatch metrics
You can view AWS Batch job metrics in the AWS Batch console, or in the CloudWatch console under
the AWS/Batch namespace.
Important
Displaying metrics in the console incurs standard CloudWatch charges. You can turn off
metric viewing at any time to stop the charges. For more information, see Amazon CloudWatch Pricing
To view metrics on the Batch metrics dashboard
-
Open the AWS Batch console
. -
In the navigation pane, choose Dashboard, then choose the Batch metrics tab.
-
Turn on Display CloudWatch metrics.
-
Under Filter metrics, select up to five Job queues and the Metrics you want to graph.
-
Choose Apply filters. The console shows one graph per metric, with one line per job queue.
-
To start over, choose Clear filters to remove your selections, or Reset filters to restore the defaults.
-
(Optional) To reuse a combination of job queues and metrics, work with a saved filter set. Filter sets are saved to your AWS user preferences, so they persist across sessions, and you can save up to five. Manage them from the Saved filter sets dropdown and the arrow menu on the Clear filters button, then choose any of the following:
-
Save - Select your job queues and metrics, then choose Save as new filter set. Enter a unique name and optionally mark it as the default.
-
Apply - Choose a set from the Saved filter sets dropdown, then choose Apply filters.
-
Update - With a set selected, change your selections and choose Update current filter set.
-
Set as default - With a set selected, choose Settings, then Set as default. The default set loads automatically when you open the tab.
-
Delete - With a set selected, choose Delete current filter set.
-
To view metrics for a single job queue
-
In the navigation pane, choose Job queues, then choose a job queue.
-
Choose the Monitoring tab.
-
Turn on Display CloudWatch metrics.
-
In the State transition metrics, Lifecycle event metrics, and Duration metrics sections, use each metrics selector to choose which metrics to graph.
The metrics shown are scoped to this job queue. Your selections are saved per job queue and persist across sessions.
To work with these metrics directly in CloudWatch — for example, to build dashboards or
set alarms — open the CloudWatch consoleAWS/Batch namespace.