Content Domain 3: Deployment and Orchestration of ML and AI Workflows
Task 3.1: Manage deployment infrastructure for ML and AI model types.
Skill 3.1.1: Select appropriate compute environments and deployment targets.
Skill 3.1.2: Select deployment orchestrators and multi-model or multi-container deployment strategies.
Skill 3.1.3: Select model inference strategies (for example, real-time and batch processing).
Skill 3.1.4: Evaluate and select appropriate foundation model (FM) deployment options.
Skill 3.1.5: Deploy models that were built outside of AWS into AWS environments (for example, Amazon SageMaker AI, Amazon Bedrock Custom Model Import).
Skill 3.1.6: Deploy and configure agents for specific tasks, integration with other services and tools, and agent communication protocols.
Skill 3.1.7: Configure FM deployment, model hosting, and resource allocation.
Skill 3.1.8: Apply Retrieval Augmented Generation (RAG) system configurations (for example, retrieval strategies, reranking).
Task 3.2: Provision and configure resources for ML and AI workloads based on existing architecture and requirements.
Skill 3.2.1: Optimize resource provisioning between on-demand and provisioned resources for performance and cost efficiency.
Skill 3.2.2: Automate compute resource provisioning with integrated communication between stacks and orchestration services.
Skill 3.2.3: Build and maintain containers for ML and AI workloads.
Skill 3.2.4: Configure SageMaker AI endpoints within VPC network environments.
Skill 3.2.5: Deploy and host models programmatically (for example, by using the SageMaker AI SDK for Python, AWS CLI, Boto3).
Skill 3.2.6: Select specific metrics for auto scaling implementations.
Skill 3.2.7: Create and manage Amazon Bedrock knowledge bases with vector database configurations, document indexing, and retrieval optimization.
Skill 3.2.8: Implement retrieval pipelines to meet business needs.
Skill 3.2.9: Implement agent state management systems.
Skill 3.2.10: Implement AI-specific resource scaling for GPU workloads.
Skill 3.2.11: Deploy agentic workflow infrastructure.
Task 3.3: Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloads.
Skill 3.3.1: Implement automated deployment strategies and rollback actions.
Skill 3.3.2: Configure and troubleshoot AWS CodeBuild, AWS CodeCommit, AWS CodeDeploy, AWS CodePipeline, and AWS CodeConnections.
Skill 3.3.3: Configure training and inference jobs.
Skill 3.3.4: Configure automated testing strategies within CI/CD pipelines for traditional ML and AI workloads.
Skill 3.3.5: Build and integrate mechanisms to re-train models.
Skill 3.3.6: Manage model versions for repeatability and audits (for example, SageMaker Model Registry, MLflow on SageMaker AI).
Skill 3.3.7: Manage prompts (for example, Amazon Bedrock Prompt Management).
Skill 3.3.8: Implement automated agent deployment pipelines and agent version management.
Skill 3.3.9: Implement AI model testing frameworks, including prompt testing.
Skill 3.3.10: Configure FM deployment automation with fine-tuned model versioning.
Skill 3.3.11: Configure AI-specific pipeline orchestration for RAG system updates and knowledge base refresh cycles.