FM Evaluation and Performance
This topic focuses on human evaluation, benchmarks, Bedrock Model Evaluation, ROUGE, BLEU, BERTScore, LLM-as-a-judge, RAG evaluation, agents, workflows, and business metrics.
How to study for AWS Certified AI Practitioner
Treat each question as a business and governance decision: identify the AI pattern, choose the right AWS capability, then add responsible AI, cost, security, and evaluation controls.
Core concepts
Concept 1
FM evaluation should combine technical metrics, human review, business outcomes, safety checks, and task-specific performance.
Exam cue: Use human-in-the-loop evaluation when judgment, safety, or subjective quality matters.
Concept 2
Text generation metrics such as ROUGE, BLEU, and BERTScore measure different kinds of similarity and should be interpreted carefully.
Exam cue: Use business metrics when the question asks whether the application meets objectives.
Concept 3
Applications built with RAG, agents, and workflows need evaluation of retrieval quality, task completion, source grounding, and user satisfaction.
Exam cue: Evaluate both the model and the surrounding application workflow.
Risk pitfalls and guardrails
Treating one benchmark as proof the application is ready.
Guardrail: Avoid choosing GenAI because it sounds modern, trusting fluent output without validation, or ignoring privacy, cost, and governance requirements.
Using a text similarity metric when the business goal is task completion.
Guardrail: Avoid choosing GenAI because it sounds modern, trusting fluent output without validation, or ignoring privacy, cost, and governance requirements.
Ignoring retrieval quality in a RAG application.
Guardrail: Avoid choosing GenAI because it sounds modern, trusting fluent output without validation, or ignoring privacy, cost, and governance requirements.
Memory anchors
Human-in-the-Loop
Human-in-the-loop evaluation uses people to review outputs, risks, or decisions.
Benchmark Dataset
A benchmark dataset provides standard examples for comparing model behavior.
ROUGE
ROUGE is commonly used to compare generated summaries with reference summaries.
BLEU
BLEU is commonly used to compare generated text with reference translations or text.
BERTScore
BERTScore uses embeddings to compare semantic similarity between texts.
LLM-as-a-Judge
LLM-as-a-judge uses a model to evaluate outputs against criteria.
Task Completion Rate
Task completion rate measures how often users successfully finish the intended task.
Cost per Interaction
Cost per interaction measures the expense of serving each user exchange or task.
Checkpoint rule
Do the check-up only after you can summarize each concept in one sentence and identify one dangerous pitfall from memory.
Knowledge Check (after reading)
Short check-up to confirm understanding of this module.
Check-up Questions
Before comparing foundation models for a support application, what should the team define first?
Which evaluation is especially useful for judging whether generated explanations are clear and helpful to the intended audience?
Answer all questions to submit.
Next step personalized recommendations
Continue learning
Move forward only after this module is stable.
What is Pass Harbor?
Completely free exam prep for 317 U.S. exams.
- Practice questions
- Flashcards
- Study guides
- Mock exams
- No registration
- No paywall
- Start instantly
“No more expensive exam prep. Quality study tools should be accessible to everyone.”
