View a markdown version of this page

Content Domain 4: Operating, Monitoring, and Securing ML and AI Solutions - AWS Certified Machine Learning Engineer - Associate

Content Domain 4: Operating, Monitoring, and Securing ML and AI Solutions

Task 4.1: Monitor ML and AI model inference and performance.

  • Skill 4.1.1: Monitor model performance in production by using Amazon CloudWatch generative AI observability, Amazon Bedrock Model Evaluation, and drift detection pipelines.

  • Skill 4.1.2: Monitor workflows to detect anomalies or errors in data processing or model inference.

  • Skill 4.1.3: Detect changes in data distribution that can affect model performance.

  • Skill 4.1.4: Monitor model performance in production by using A/B testing.

  • Skill 4.1.5: Monitor and automate the management of agent performance and coordination (for example, coordination failure detection, truncated streaming, tool failures).

  • Skill 4.1.6: Configure AI-specific performance monitoring for foundation models (FMs), such as Amazon Bedrock evaluations.

Task 4.2: Optimize and manage ML and AI infrastructure costs and performance.

  • Skill 4.2.1: Select inference instance families to optimize performance and cost.

  • Skill 4.2.2: Configure and use tools to troubleshoot and analyze resources (for example, Amazon CloudWatch, Amazon Bedrock AgentCore Observability, AWS X-Ray).

  • Skill 4.2.3: Set up dashboards to monitor performance metrics.

  • Skill 4.2.4: Optimize capacity for cost, performance, and reliability.

  • Skill 4.2.5: Optimize costs and set cost quotas by using appropriate cost management tools.

  • Skill 4.2.6: Optimize infrastructure costs by selecting purchasing options.

  • Skill 4.2.7: Evaluate cost implications of using FMs for inference in production.

  • Skill 4.2.8: Monitor agent resource consumption patterns.

  • Skill 4.2.9: Manage FM inference costs with usage optimization.

  • Skill 4.2.10: Monitor AI-specific cost patterns (for example, token usage optimization, embedding computation costs, vector database storage optimization).

Task 4.3: Secure ML and AI workloads and model endpoints.

  • Skill 4.3.1: Secure continuous integration and continuous delivery (CI/CD) pipelines by checking for code and image vulnerabilities (for example, by using Amazon CodeGuru, Amazon Inspector).

  • Skill 4.3.2: Configure least privilege access to ML and AI artifacts.

  • Skill 4.3.3: Configure IAM policies and roles for users and applications in ML and AI systems.

  • Skill 4.3.4: Configure comprehensive monitoring, auditing, compliance, and logging for ML and AI systems (for example, AWS CloudTrail, AWS Config).

  • Skill 4.3.5: Troubleshoot and debug security issues in ML and AI systems.

  • Skill 4.3.6: Create VPCs, subnets, and security groups to securely isolate ML and AI systems.

  • Skill 4.3.7: Identify and mitigate security risks and vulnerabilities in ML and AI systems.

  • Skill 4.3.8: Select the appropriate credential type to access FMs (for example, Amazon Bedrock API keys, IAM credentials).

  • Skill 4.3.9: Implement safeguards and sensitive data protection to meet application requirements and responsible AI policies (for example, by using Amazon Bedrock Guardrails).