This prompt is designed for QA engineers and platform teams who have already migrated active AI sessions to new instruction versions and need an automated gate to validate the results. The core job-to-be-done is comparing pre-migration and post-migration outputs against a set of defined acceptance criteria and flagging any regressions that should block the rollout. It assumes you have recorded pre-migration behavior, a batch of migrated sessions, and a clear set of criteria that define what 'equivalent or better' means for your use case. Without these prerequisites, the prompt cannot produce a reliable verification report.
Prompt
Migration Success Verification Prompt for QA Pipelines

When to Use This Prompt
Use this prompt to verify that migrated AI sessions produce behavior equivalent to or better than pre-migration behavior before a rollout reaches production.
The ideal user is a QA engineer or release manager running this prompt inside a CI/CD pipeline's QA gate, triggered after a session migration dry-run or batch migration job completes. Required context includes the pre-migration session transcript, the post-migration session transcript, the acceptance criteria document, and any domain-specific rules that define acceptable behavioral variance. Do not use this prompt for initial prompt design validation, for testing sessions that were never migrated, or for evaluating raw model outputs without a migration context. It is specifically scoped to regression detection across an instruction version boundary.
In practice, you will wire this prompt into an automated test harness that feeds paired session transcripts and acceptance criteria, then parses the structured verification report. The report should include a pass/fail determination per session, a regression severity classification for any failures, and a summary of which acceptance criteria were violated. Common failure modes include false positives from acceptable behavioral variance that wasn't captured in the criteria, and false negatives from criteria that are too permissive. To mitigate these, maintain a living set of acceptance criteria that evolves with your instruction changes, and always include a human review step for any session flagged as a regression before blocking a production rollout.
Use Case Fit
Where the Migration Success Verification Prompt works, where it fails, and what you must provide before running it in a QA pipeline.
Good Fit: Pre- and Post-Migration Snapshots Exist
Use when: you have captured structured session state snapshots before and after an instruction migration. The prompt compares behavior against acceptance criteria and flags regressions. Guardrail: require both snapshots to pass a completeness check before verification runs.
Bad Fit: No Baseline Behavior Recorded
Avoid when: only post-migration traces are available with no pre-migration reference. The prompt cannot verify correctness without a known-good baseline. Guardrail: gate verification on the presence of a validated pre-migration snapshot; otherwise route to manual QA.
Required Inputs: Acceptance Criteria and Session Pairs
What you need: a set of acceptance criteria, paired pre- and post-migration session traces, and the instruction version identifiers for both sides. Guardrail: validate that acceptance criteria are specific, testable, and mapped to observable behaviors before feeding them into the prompt.
Operational Risk: Silent Behavioral Drift
What to watch: the prompt may report a pass when subtle behavioral changes exist outside the acceptance criteria. Guardrail: supplement automated verification with a manual spot-check on a sample of migrated sessions, especially for tone, refusal style, and edge-case handling.
Operational Risk: Criteria Coverage Gaps
What to watch: acceptance criteria that cover happy paths but miss edge cases, tool-call ordering, or error-recovery behavior. Guardrail: run a coverage audit on acceptance criteria before migration, and flag any session behavior category with zero criteria.
Bad Fit: Unstable or Non-Deterministic Sessions
Avoid when: session behavior varies significantly across runs due to temperature, tool latency, or model non-determinism. The prompt may flag false regressions. Guardrail: run multiple pre-migration samples and establish a behavior envelope; only compare post-migration runs that fall within expected variance.
Copy-Ready Prompt Template
A reusable prompt template for verifying that migrated sessions behave correctly after an instruction update, producing a structured QA report.
This template is the core of the Migration Success Verification Prompt for QA Pipelines. It is designed to be copied directly into your testing harness, an automated eval runner, or a manual QA review interface. The prompt instructs the model to act as a QA engineer, comparing pre- and post-migration behavior for a given session against a set of acceptance criteria. The output is a structured verification report that flags regressions, confirms expected changes, and notes any ambiguous behavior.
textYou are a QA engineer validating a session migration. Your task is to compare the behavior of an AI agent before and after an instruction update to ensure the migration was successful. You will be given the following inputs: - [PRE_MIGRATION_TRANSCRIPT]: The full conversation transcript before the instruction update. - [POST_MIGRATION_TRANSCRIPT]: The full conversation transcript after the instruction update, starting from the same user context. - [ACCEPTANCE_CRITERIA]: A list of specific, verifiable behavioral requirements that the post-migration session must meet. - [MIGRATION_CONTEXT]: A description of what changed in the instructions and why. Your task is to produce a Verification Report in the following JSON format: { "migration_id": "string", "overall_status": "PASS" | "FAIL" | "INCONCLUSIVE", "criteria_results": [ { "criterion_id": "string", "description": "string", "status": "MET" | "NOT_MET" | "UNABLE_TO_VERIFY", "evidence": "A direct quote or specific reference from the transcripts.", "rationale": "A brief explanation of why the status was assigned." } ], "regressions_found": [ { "description": "string", "pre_migration_evidence": "string", "post_migration_evidence": "string", "severity": "CRITICAL" | "HIGH" | "MEDIUM" | "LOW" } ], "unexpected_changes": [ "A description of any behavioral change not covered by the acceptance criteria." ], "summary": "A concise 2-3 sentence summary of the migration outcome." } Follow these constraints: - [CONSTRAINTS]: Base all findings strictly on the provided transcripts. Do not infer intent. If evidence is insufficient, use "UNABLE_TO_VERIFY". A regression is any behavior that was correct pre-migration but is incorrect or missing post-migration.
To adapt this template, replace the square-bracket placeholders with your test data. For automated pipelines, [PRE_MIGRATION_TRANSCRIPT] and [POST_MIGRATION_TRANSCRIPT] can be loaded from your test case database. The [ACCEPTANCE_CRITERIA] should be a structured list derived from your migration plan's success metrics. The [CONSTRAINTS] placeholder can be extended with domain-specific rules, such as requiring citations to specific document sections or flagging tone violations. After running this prompt, always validate the output JSON against your expected schema before marking a migration as verified. For high-risk migrations, the output of this prompt should be a mandatory input to a human approval gate, not an automated pass/fail decision.
Prompt Variables
Inputs the Migration Success Verification Prompt needs to produce a reliable QA report. Validate each placeholder before running the prompt to prevent false-pass or false-fail signals in your regression pipeline.
| Placeholder | Purpose | Example | Validation Notes |
|---|---|---|---|
[PRE_MIGRATION_TRACE] | The agent's behavior log from a session running the old instruction version | {"session_id": "sess-982", "turns": [{"role": "user", "content": "..."}, {"role": "assistant", "tool_calls": [...]}]} | Must be valid JSON with at least one complete turn. Reject if trace is truncated, missing tool_call responses, or contains unresolved placeholders. |
[POST_MIGRATION_TRACE] | The agent's behavior log from the same session after migration to the new instruction version | {"session_id": "sess-982", "turns": [{"role": "assistant", "content": "..."}]} | Must be valid JSON. Compare session_id against [PRE_MIGRATION_TRACE]; reject if IDs don't match or if trace contains only system-level events with no assistant turns. |
[ACCEPTANCE_CRITERIA] | The pass/fail rules defining acceptable post-migration behavior | ["No tool call signature changes for unchanged tools", "Response latency delta < 500ms", "Refusal rate does not increase by >5%"] | Must be a non-empty array of strings. Reject if any criterion is ambiguous (e.g., 'behaves correctly') without a measurable threshold. Each criterion must be parseable into an automated check. |
[INSTRUCTION_DIFF] | A semantic diff between old and new instruction versions, not a raw text diff | {"added_rules": ["Rule 14: Escalate billing disputes to Tier 2"], "removed_rules": ["Rule 9: Auto-close after 3 turns"], "modified_rules": [{"old": "...", "new": "..."}]} | Must be valid JSON with at least one of added_rules, removed_rules, or modified_rules non-empty. Reject if diff is a raw text patch without semantic categorization. |
[REGRESSION_TEST_CASES] | Known edge cases that must produce identical or explicitly changed behavior post-migration | [{"input": "I want a refund", "expected_pre": "refund_initiated", "expected_post": "refund_initiated"}, {"input": "Cancel my account", "expected_pre": "escalate_to_retention", "expected_post": "escalate_to_retention"}] | Must be a non-empty array of objects with input, expected_pre, and expected_post fields. Reject if any expected_post value is null without an explicit 'behavior_removed' flag. |
[MIGRATION_METADATA] | Version identifiers and rollout context for the migration under test | {"old_version": "v2.4.1", "new_version": "v3.0.0", "migration_batch": "canary-3", "rollout_timestamp": "2025-06-15T14:30:00Z"} | Must include old_version and new_version as non-empty strings. Reject if version strings are identical or if rollout_timestamp is in the future. Batch identifier required for canary or staged rollouts. |
[TOOL_CONTRACT_SNAPSHOT] | The tool definitions and schemas available to the agent in both versions | {"old_tools": [{"name": "refund", "parameters": {...}}], "new_tools": [{"name": "refund", "parameters": {...}}]} | Must be valid JSON with old_tools and new_tools arrays. Reject if any tool name appears in one version but not the other without an explicit deprecation or addition note in [INSTRUCTION_DIFF]. |
Implementation Harness Notes
How to wire the Migration Success Verification Prompt into a QA pipeline or CI/CD workflow.
The Migration Success Verification Prompt is designed to be a gate in your CI/CD pipeline, not a one-off manual check. It should run automatically after every instruction version bump or session migration event. The prompt expects a structured input payload containing pre-migration and post-migration session traces, the acceptance criteria for the migration, and the version identifiers for both the old and new instruction sets. The output is a structured verification report that your pipeline can parse to decide whether to promote the migration, flag regressions, or block the rollout.
To integrate this prompt, wrap it in a harness that first collects the required inputs: [PRE_MIGRATION_TRACE], [POST_MIGRATION_TRACE], [ACCEPTANCE_CRITERIA], [OLD_INSTRUCTION_VERSION], and [NEW_INSTRUCTION_VERSION]. The harness should validate that these inputs are non-empty and that the version identifiers match the expected migration pair. After calling the model, parse the [OUTPUT_SCHEMA]—a JSON object containing a verification_status (PASS, FAIL, or INCONCLUSIVE), a list of regression_flags, and a behavioral_comparison_summary. Implement a retry loop with a maximum of 2 retries if the model returns malformed JSON or fails to include the required fields. Log every attempt, including the raw model response and the parsed output, for auditability. For high-risk migrations in regulated environments, add a human approval step: if verification_status is not PASS, or if any regression_flags have a severity of HIGH, block the pipeline and route the report to a review queue before proceeding.
Choose a model with strong instruction-following and structured output capabilities for this task. The prompt relies on precise comparison and schema adherence, so smaller or less capable models may produce inconsistent verification reports. If you are running this in a CI/CD environment with cost constraints, consider caching the prompt prefix that contains the static instructions and output schema. The dynamic inputs—the traces and criteria—should be appended after the cached prefix. Before deploying this harness to production, build a golden test set of known migration scenarios: one where the migration should PASS, one where it should FAIL due to a behavioral regression, and one where the session state was corrupted. Run these through the harness and assert that the verification_status matches the expected outcome for each case. This pre-release eval gate will catch prompt brittleness before it affects a real rollout.
Expected Output Contract
Fields, format, and validation rules for the migration success verification report. Use this contract to build downstream parsers, dashboards, and automated pass/fail gates.
| Field or Element | Type or Format | Required | Validation Rule |
|---|---|---|---|
report_id | string (UUID v4) | Must parse as valid UUID v4. Reject if missing or malformed. | |
migration_id | string | Must match the [MIGRATION_ID] input exactly. Case-sensitive match required. | |
session_id | string | Must match the [SESSION_ID] input exactly. Case-sensitive match required. | |
verdict | enum: pass | fail | partial | blocked | Must be one of the four allowed values. Reject any other string. | |
acceptance_criteria_results | array of objects | Each object must contain criterion_id (string), met (boolean), and evidence (string, max 500 chars). Array must not be empty. | |
regression_flags | array of objects | Each object must contain test_name (string), severity (enum: low | medium | high | critical), and description (string). Empty array allowed only if verdict is pass. | |
behavior_comparison_summary | string | Must be non-empty. Max 2000 characters. Must reference at least one pre-migration and one post-migration behavior observation. | |
human_review_required | boolean | Must be true if verdict is fail, partial, or blocked, or if any regression_flag severity is high or critical. Validate with boolean type check, not string coercion. |
Common Failure Modes
What breaks first when verifying migration success in QA pipelines and how to guard against it.
Behavioral Equivalence Blind Spots
What to watch: The verification prompt confirms surface-level output format matches but misses subtle behavioral regressions—tone shifts, refusal pattern changes, or degraded reasoning on edge cases that weren't in the test suite. Guardrail: Include behavioral probes in the verification criteria that test refusal boundaries, ambiguity handling, and multi-turn consistency, not just output schema conformance.
Stale Acceptance Criteria Drift
What to watch: The acceptance criteria used for verification were written against the old instruction version and haven't been updated to reflect new capabilities, deprecated behaviors, or intentional breaking changes. The prompt passes verification but validates against wrong expectations. Guardrail: Version-lock acceptance criteria to instruction versions and require explicit criteria review as part of the migration verification step before running comparisons.
Session Context Contamination
What to watch: Migrated sessions carry conversation history, tool outputs, or user corrections from the old instruction regime that conflict with new rules. The verification prompt evaluates post-migration behavior in isolation but misses contamination from pre-migration context. Guardrail: Include pre-migration context replay in verification tests and flag outputs where old context appears to override new instructions.
False-Positive Regression Flags
What to watch: The verification prompt flags intentional behavior changes as regressions because the diff threshold is too sensitive or the acceptance criteria don't distinguish between breaking changes and deliberate improvements. Teams waste cycles investigating non-issues. Guardrail: Add an intentional-change allowlist to the verification prompt and require explicit justification when a flagged difference matches a documented behavioral change in the migration changelog.
Incomplete Session Coverage Sampling
What to watch: The verification prompt only tests a handful of happy-path sessions and misses failure modes that surface in long-running conversations, sessions with tool errors, or conversations with adversarial user inputs. Guardrail: Stratify the verification sample by session length, tool-use density, error-recovery count, and user correction frequency. Require minimum coverage thresholds per stratum before the verification report is considered complete.
Silent Instruction Priority Inversion
What to watch: After migration, new system instructions conflict with residual user-level directives or tool-output context from the pre-migration session, and the model silently resolves the conflict incorrectly without surfacing ambiguity. The verification prompt doesn't detect this because it doesn't probe instruction priority. Guardrail: Include explicit instruction-conflict probes in the verification suite that test whether system-level migration rules correctly override stale user or tool context from the old session.
Evaluation Rubric
Use this rubric to test the Migration Success Verification Prompt before shipping it to production. Each criterion targets a specific failure mode observed in QA pipeline migrations.
| Criterion | Pass Standard | Failure Signal | Test Method |
|---|---|---|---|
Pre/Post Behavior Comparison | Every acceptance criterion is paired with a pre-migration and post-migration observation that references specific evidence from the session trace. | Criteria listed without paired observations, or observations that paraphrase the criterion without citing session evidence. | Parse output for [ACCEPTANCE_CRITERIA] array. Assert each item has non-null pre_migration_observation and post_migration_observation fields with distinct trace references. |
Regression Flagging | Any post-migration behavior that violates a previously passing criterion is flagged with severity, affected criterion ID, and a reproduction note. | Regressions present in the trace but missing from the flags array, or flags without severity and criterion ID linkage. | Inject a known regression into the test fixture. Assert output contains a flag with severity >= 'medium', correct criterion_id, and a non-empty reproduction_note. |
Session State Preservation | Verification report confirms that conversation context, pending actions, and user intent survived migration without loss or corruption. | Report claims state preservation but omits checks for pending actions or user intent, or reports state loss without flagging it as a regression. | Provide a fixture with a pending action and explicit user intent. Assert state_preservation section lists both and marks each as 'preserved' or flags loss. |
Instruction Version Attribution | Every behavioral check cites which instruction version was active when the behavior occurred. | Behavioral observations that reference instructions without version identifiers, or version references that don't match the migration timeline. | Assert each observation includes an instruction_version field matching the expected pre-migration or post-migration version string from the fixture. |
Edge-Case Coverage | Report includes checks for session boundary conditions: empty context, maximum-length sessions, interrupted tool calls, and multi-role transitions. | Report only covers happy-path criteria and ignores edge cases present in the test fixture. | Use a fixture with an interrupted tool call and a role transition. Assert edge_case_checks array contains entries for both scenarios with pass or flag outcomes. |
Confidence and Uncertainty | Low-confidence findings are explicitly marked with a confidence score and a recommendation for human review. | Report presents all findings with equal certainty, or low-confidence findings lack a human review recommendation. | Inject ambiguous migration behavior. Assert at least one finding has confidence_score < 0.8 and human_review_recommended equals true. |
Output Schema Compliance | Report matches the [OUTPUT_SCHEMA] exactly: all required fields present, no extra top-level keys, enum values within allowed sets. | Missing required fields, extra keys, or severity values outside the allowed enum. | Validate output against the declared JSON schema. Assert no schema violations. Reject if required fields are null unless explicitly allowed. |
Enabling Efficiency, Speed & Accuracy
Intelligent Analysis, Decision & Execution
We build AI systems for teams that need search across company data, workflow automation across tools, or AI features inside products and internal software.
Talk to Us
Search across company data
Give teams answers from docs, tickets, runbooks, and product data with sources and permissions.
Useful when people spend too long searching or get different answers from different systems.

Automate internal workflows
Use AI to route work, draft outputs, trigger actions, and keep approvals and logs in place.
Useful when repetitive work moves across multiple tools and teams.

Add AI to products and internal tools
Build assistants, guided actions, or decision support into the software your team or customers already use.
Useful when AI needs to be part of the product, not a separate tool.
Adapt This Prompt
How to adapt
Add strict JSON schema validation on the verification report output. Wrap the prompt in a retry loop with exponential backoff for malformed responses. Log every verification result with session ID, instruction version pair, and timestamp. Integrate with your CI/CD pipeline so migration verification blocks deployment if regression flags exceed threshold.
Prompt modification
Add to [CONSTRAINTS]: "Return ONLY valid JSON matching the [OUTPUT_SCHEMA]. If you cannot determine a field value, use null, not a placeholder string." Add a [SESSION_METADATA] input block containing session_id, pre_migration_version, post_migration_version, and migration_timestamp for traceability.
Watch for
- Silent format drift when model outputs valid JSON but wrong field names
- Missing human review gate for high-severity regressions
- Retry loops masking persistent failures instead of surfacing them

About the author
Prasad Kumkar
CEO & MD, Inference Systems
Prasad Kumkar is the CEO & MD of Inference Systems and writes about AI systems architecture, LLM infrastructure, model serving, evaluation, and production deployment. Over 5+ years, he has worked across computer vision models, L5 autonomous vehicle systems, and LLM research, with a focus on taking complex AI ideas into real-world engineering systems.
His work and writing cover AI systems, large language models, AI agents, multimodal systems, autonomous systems, inference optimization, RAG, evaluation, and production AI engineering.
Partnered with leading AI, data, and software stack.
How We Work
Custom AI workflows for your Business
One-fit-all AI don't work for modern businesses. At Inferensys, we aim to understand your business & custom requirements; which we use to define most efficient agentic workflows, the data, and the tools for your business.
01
Review the use case
We understand the task, the users, and where AI can actually help.
Read more02
Pick the right approach
We define what needs search, automation, or product integration.
Read more03
Build the first useful version
We implement the part that proves the value first.
Read more04
Improve from there
We add the checks and visibility needed to keep it useful.
Read moreThe first call is a practical review of your use case and the right next step.
Talk to Us