Use this prompt when a user challenges a RAG-augmented assistant's response by providing their own counter-evidence, such as a document, link, or specific data point. The core job-to-be-done is not simply accepting the correction or re-retrieving, but performing a structured comparison between the user's provided evidence and the assistant's original sources. The assistant must assess which evidence is stronger based on criteria like recency, authority, and direct relevance, then produce a corrected response that explicitly explains the resolution and updates citations accordingly. The ideal user is a developer building a high-stakes assistant for research, legal, or compliance workflows where source trustworthiness is paramount and users are domain experts who can supply better information.
Prompt
Correction with New Evidence Prompt Template

When to Use This Prompt
Defines the specific job, required context, and boundaries for the Correction with New Evidence prompt.
This prompt is essential when the cost of an incorrect resolution is high. For example, a financial analyst might correct a market summary by providing a link to a more recent SEC filing, or a legal researcher might supply a specific case that overrules the assistant's cited precedent. In these cases, the assistant must not blindly trust the user or stubbornly defend its retrieval. Instead, it must execute a transparent evidence evaluation step. The required context includes the assistant's original full response with citations, the user's correction message, and the user's new evidence. Without all three inputs, the comparison logic will fail. The prompt template includes placeholders for [ORIGINAL_RESPONSE], [USER_CORRECTION_MESSAGE], and [USER_EVIDENCE] to enforce this.
Do not use this prompt for simple factual corrections where no new evidence is provided. If a user says 'that date is wrong' without offering a source, a different prompt for claim reversal or state rollback is more appropriate. Similarly, avoid this prompt if the correction only requires re-retrieving from the same knowledge base with a better query. This prompt is specifically for the conflict-resolution scenario where two competing sources exist and a judgment call is required. Before implementing, ensure you have a logging mechanism to capture these evidence-comparison decisions, as they create a critical audit trail for understanding why the assistant changed its position.
Use Case Fit
Where the Correction with New Evidence prompt works reliably and where it introduces unacceptable risk. Use this card grid to decide whether to deploy this prompt, add a human review step, or choose a different correction strategy.
Good Fit: RAG Systems with Retrievable Sources
Use when: the assistant's original claim was grounded in a specific, retrievable source, and the user provides a counter-source or specific factual contradiction. The prompt excels at comparing evidence sets and updating citations. Guardrail: always re-retrieve the original source and the user's cited evidence before running the prompt to prevent hallucinated source comparisons.
Bad Fit: Subjective or Opinion-Based Corrections
Avoid when: the user's correction is a matter of interpretation, style preference, or opinion rather than a verifiable factual claim. The evidence-weighting logic will fabricate false confidence scores for subjective disputes. Guardrail: route opinion-based corrections to a separate clarification prompt; only use this template when both sides can cite a source.
Required Inputs: Original Source and User Evidence
Risk: running this prompt without the assistant's original retrieved source and the user's counter-evidence forces the model to invent both, producing a confident but fabricated resolution. Guardrail: the harness must inject [ORIGINAL_SOURCE], [ORIGINAL_CLAIM], and [USER_EVIDENCE] as explicit, verbatim inputs. If any are missing, abort and request clarification instead.
Operational Risk: Evidence Quality Asymmetry
Risk: the model may defer to the user's evidence simply because it is newer in the context window, or reject it because the original source appears more authoritative based on surface features like domain name. Guardrail: include an explicit evidence quality assessment step in the prompt that evaluates recency, authority, and relevance for both sources before declaring a winner. Log this assessment for audit.
Operational Risk: Correction Cascades in Multi-Turn Sessions
Risk: accepting a user's evidence and updating one claim can invalidate downstream conclusions, tool calls, or state updates that depended on the original claim. Guardrail: pair this prompt with a state rollback or dependency check step. After producing the corrected response, flag any downstream turns that need re-evaluation rather than silently continuing.
Bad Fit: Real-Time or High-Latency-Sensitive Workflows
Avoid when: the correction must be resolved in under one second. This prompt requires re-retrieval, evidence comparison, and structured output generation, which adds latency. Guardrail: for real-time systems, use a lightweight correction detection prompt first. If evidence comparison is needed, acknowledge the correction immediately and process the full evidence review asynchronously, updating the response when complete.
Copy-Ready Prompt Template
A reusable system prompt for evaluating user-provided counter-evidence against original sources and producing a corrected, re-cited response.
This template is designed for RAG-augmented assistants where a user challenges a previous answer with new evidence. The prompt instructs the model to act as an evidence arbiter, not just a polite apologizer. It forces a structured comparison between the assistant's original sources and the user's new information, determines which evidence prevails based on quality criteria, and generates a corrected response with updated citations. Use this as your system prompt or as a dedicated correction-handling turn instruction.
codeYou are an evidence-arbitration assistant. Your task is to re-evaluate a prior answer after a user provides counter-evidence. ## ORIGINAL CONTEXT - Original User Question: [ORIGINAL_QUESTION] - Assistant's Prior Answer: [PRIOR_ANSWER] - Original Source Citations: [ORIGINAL_SOURCES] ## USER CORRECTION - User's Correction Statement: [USER_CORRECTION] - User-Provided Evidence: [USER_EVIDENCE] ## EVIDENCE EVALUATION RULES 1. Compare the user's evidence against each original source for relevance, recency, authority, and specificity. 2. Classify the conflict for each disputed claim as one of: - `user_evidence_prevails`: User's source is more authoritative, recent, or specific. - `original_source_prevails`: Original source is stronger; user evidence is outdated, misinterpreted, or from a weaker authority. - `insufficient_to_decide`: Neither source is clearly superior; acknowledge uncertainty. 3. Do not automatically defer to the user. Evaluate evidence quality, not politeness. ## OUTPUT SCHEMA Return a JSON object with the following structure: { "conflict_assessment": [ { "disputed_claim": "string (the specific claim being challenged)", "original_source": "string (citation from original sources)", "user_evidence": "string (relevant excerpt from user's evidence)", "resolution": "user_evidence_prevails | original_source_prevails | insufficient_to_decide", "rationale": "string (brief explanation of why this resolution was chosen)" } ], "corrected_answer": "string (the full corrected response, incorporating prevailing evidence)", "updated_citations": ["string (list of citations supporting the corrected answer)"], "retracted_claims": ["string (list of original claims that are now withdrawn)"], "needs_re_retrieval": true | false, "re_retrieval_query": "string | null (suggested query if original retrieval was insufficient)" } ## CONSTRAINTS - If the user's evidence is weaker, politely explain why the original answer stands. Do not fabricate agreement. - If the original retrieval was incomplete and the user's evidence reveals a gap, set `needs_re_retrieval` to true and suggest a better query. - Never cite the user's evidence as a source unless you can verify it independently. Mark unverified user evidence clearly. - If the correction reveals a factual error in the original answer, acknowledge it directly without over-apologizing. - Preserve any parts of the original answer that are not disputed. ## RISK LEVEL [RISK_LEVEL]
To adapt this template, replace the square-bracket placeholders with your application's runtime values. [ORIGINAL_QUESTION], [PRIOR_ANSWER], and [ORIGINAL_SOURCES] should be pulled from your conversation state or session store. [USER_CORRECTION] and [USER_EVIDENCE] come from the current user turn. Set [RISK_LEVEL] to high for regulated domains to enforce stricter evidence standards and require human review for insufficient_to_decide outcomes. For low-risk applications, you can simplify the output schema to just corrected_answer and updated_citations. Before deploying, validate that your application layer can parse the JSON output, log conflict assessments for audit, and trigger re-retrieval when needs_re_retrieval is true. Test against cases where the user provides weak evidence, strong evidence, and no evidence at all to ensure the model doesn't default to blind agreement.
Prompt Variables
Inputs the Correction with New Evidence prompt needs to work reliably. Validate each before injection into the template. Placeholders marked with square brackets must be resolved at runtime.
| Placeholder | Purpose | Example | Validation Notes |
|---|---|---|---|
[ORIGINAL_ASSISTANT_CLAIM] | The specific assistant statement or claim the user is challenging | The Q3 revenue was $4.2M based on the 10-Q filing. | Must be a verbatim excerpt from the prior assistant turn. Null or truncated strings will cause ambiguous correction targets. Validate length > 10 chars. |
[ORIGINAL_SOURCES] | The retrieved evidence or citations the assistant used to make the original claim | [{"source_id": "doc_17", "excerpt": "Q3 revenue reached $4.2M...", "retrieval_score": 0.91}] | Must be a valid JSON array of source objects. Each object requires source_id, excerpt, and retrieval_score fields. Empty array allowed if original claim had no retrieval grounding. |
[USER_CORRECTION_TEXT] | The user's verbatim correction message including any counter-evidence they provided | That's wrong. The 10-Q was restated in November. The corrected figure is $3.8M. Here's the amended filing link. | Must be the raw user turn text. Do not pre-process or summarize. Validate non-empty string. Multi-sentence corrections are expected and valid. |
[USER_PROVIDED_EVIDENCE] | Any evidence, links, documents, or data the user included with their correction | {"links": ["https://example.com/amended-10q.pdf"], "claims": ["Corrected revenue is $3.8M"], "document_excerpts": null} | Must be a structured JSON object with links, claims, and document_excerpts arrays. All arrays can be empty. Links must pass URL format validation. Null allowed if user provided no evidence. |
[CONVERSATION_CONTEXT] | The prior conversation turns leading up to the correction for context on how the claim was introduced | [{"role": "user", "content": "What was Q3 revenue?"}, {"role": "assistant", "content": "Q3 revenue was $4.2M based on the 10-Q."}] | Must be a valid JSON array of turn objects with role and content fields. Minimum 2 turns required. Validate role values are only user or assistant. Truncate to last 10 turns if context budget is tight. |
[EVIDENCE_QUALITY_THRESHOLD] | Confidence threshold for accepting user evidence over original sources | 0.7 | Must be a float between 0.0 and 1.0. Default 0.7. Lower values make the assistant more deferential to user corrections. Validate type and range before injection. Null not allowed; use default if unspecified. |
[OUTPUT_SCHEMA] | The expected structure for the corrected response including conflict resolution fields | {"correction_accepted": true, "prevailing_evidence": "user", "corrected_claim": "...", "updated_citations": [...], "conflict_analysis": "...", "confidence": 0.92} | Must be a valid JSON Schema or example object. Validate parseable JSON. Required fields: correction_accepted (bool), prevailing_evidence (enum: user|original|insufficient), corrected_claim (string), updated_citations (array), conflict_analysis (string), confidence (float). |
Implementation Harness Notes
How to wire the Correction with New Evidence prompt into a RAG application with validation, retries, and source conflict resolution.
The Correction with New Evidence prompt operates inside a retrieval-augmented generation loop. When a user correction arrives with counter-evidence—a link, a document snippet, or a specific factual claim—the harness must first extract that evidence, then invoke this prompt with both the original retrieved sources and the user-provided evidence. The prompt's job is to evaluate which evidence prevails and produce a corrected response with updated citations. The harness is responsible for feeding the right inputs, validating the output, and deciding whether to re-retrieve, escalate, or commit the correction to session state.
Wire the prompt into your application as a dedicated correction handler that fires when the User Correction Detection prompt flags a turn as containing counter-evidence. Before calling this prompt, extract the user's evidence payload: if the user pasted a URL, fetch and chunk it; if they quoted text, capture the exact string; if they referenced a document already in your knowledge base, retrieve it fresh. Pass the original assistant response, its source citations, and the user's evidence into the [ORIGINAL_RESPONSE], [ORIGINAL_SOURCES], and [USER_EVIDENCE] placeholders respectively. On the output side, validate that the response contains: (1) an evidence assessment with a clear prevailing-evidence decision, (2) a corrected response that addresses the specific disputed claim, and (3) updated citations that include the user's evidence where it prevailed. If the output is malformed or the prevailing-evidence field is missing, retry once with a stricter schema constraint. If the prompt determines that neither evidence set is sufficient, route to a clarification or human escalation path rather than forcing a low-confidence correction.
Log every correction event with the original claim, user evidence, prevailing-evidence decision, and corrected output. This audit trail feeds the Correction Audit Trail Generation prompt and enables observability into correction failure patterns. Watch for two common harness-level failure modes: first, the prompt may accept user evidence that is itself incorrect or outdated—consider adding an evidence freshness check before calling this prompt. Second, the prompt may produce a corrected response that contradicts other session state—run the output through the State Rollback After Correction prompt to identify downstream state fields that need recomputation. For high-stakes domains, always require human review before committing a correction that reverses a sourced claim, especially when the original sources are authoritative and the user evidence is unverified.
Expected Output Contract
Fields, format, and validation rules for the JSON output produced by the Correction with New Evidence prompt. Use this contract to parse, validate, and route the model's response in your application harness.
| Field or Element | Type or Format | Required | Validation Rule |
|---|---|---|---|
correction_id | string (UUID v4) | Must be a valid UUID v4 string. Generate if not present in input. | |
original_claim | string | Must match the exact claim text from [ORIGINAL_ASSISTANT_CLAIM]. Non-empty. | |
user_evidence_summary | string | Must be a non-empty summary of the evidence provided in [USER_CORRECTION_EVIDENCE]. | |
evidence_assessment | object | Must contain 'prevailing_evidence' (enum: user | assistant | inconclusive) and 'confidence' (float 0.0-1.0). | |
source_conflict_analysis | array of objects | Each object must have 'source_id', 'claim_supported', 'relevance_score' (float 0.0-1.0), and 'reliability_note'. Array must not be empty. | |
corrected_response | string | Must differ from [ORIGINAL_ASSISTANT_RESPONSE] if prevailing_evidence is 'user'. Must not introduce new claims unsupported by evidence. | |
updated_citations | array of objects | Each object must have 'citation_id', 'source_text', and 'used_in_correction' (boolean). Must include all citations referenced in corrected_response. | |
requires_human_review | boolean | Must be true if evidence_assessment.confidence < [CONFIDENCE_THRESHOLD] or prevailing_evidence is 'inconclusive'. |
Common Failure Modes
What breaks first when users provide new evidence to correct an assistant, and how to guard against cascading failures.
Source Conflict Deadlock
What to watch: The model cannot resolve a conflict between the user's new evidence and the original retrieved sources, resulting in a non-committal or contradictory response that satisfies neither side. Guardrail: Implement a strict evidence hierarchy (e.g., user-provided document > real-time API > vector store) and force a tie-breaker decision with explicit rationale rather than allowing equivocation.
Evidence Quality Blindness
What to watch: The assistant accepts user-provided evidence at face value without assessing its credibility, leading to a correction based on a hallucinated or low-quality user source. Guardrail: Add an explicit evidence quality assessment step that checks for internal consistency, source attribution, and contradiction with known facts before accepting the correction.
Partial Correction Overwrite
What to watch: A user correction targeting one specific claim causes the assistant to discard the entire prior response, including accurate and uncontested information. Guardrail: Require the prompt to isolate the disputed claim span, preserve all non-disputed content, and produce a merged response that only modifies the specific corrected segment.
Citation Drift After Correction
What to watch: The corrected response introduces new claims that lack proper citations or reuses original citations that no longer support the updated content. Guardrail: Enforce a post-correction citation verification step that maps every factual statement in the corrected response to a specific source passage, flagging unsupported claims for removal.
Correction Cascade Instability
What to watch: Accepting one piece of new evidence triggers a chain of dependent corrections across multiple prior turns, tool outputs, or state fields, causing the assistant to unravel its entire session context. Guardrail: Build a dependency graph of claims before applying the correction, identify all downstream affected items, and apply changes in a controlled order with a maximum cascade depth limit.
Over-Compliance with User Evidence
What to watch: The assistant defers too readily to user-provided evidence even when the original sources are demonstrably more authoritative, eroding the system's reliability and encouraging users to game the assistant with fabricated evidence. Guardrail: Implement a confidence threshold that compares the assistant's original source quality score against the user's evidence quality score, and escalate to human review when the conflict exceeds a defined gap.
Evaluation Rubric
Score each test case against these criteria before shipping the Correction with New Evidence prompt. Use a 0-1 scale for automated checks and a structured rubric for human review.
| Criterion | Pass Standard | Failure Signal | Test Method |
|---|---|---|---|
Evidence Comparison Completeness | Output explicitly compares user-provided evidence against each original source claim, noting agreement, contradiction, or independence. | Output accepts user evidence without comparing to original sources, or dismisses user evidence without analysis. | LLM-as-judge check: Does the output contain a comparison block mapping user evidence to each original source? |
Source Conflict Resolution | When user evidence contradicts original sources, output states which evidence prevails and provides a clear, non-evasive rationale based on recency, authority, or specificity. | Output equivocates, presents both sides without resolution, or defaults to original sources without justification. | Schema check: |
Evidence Quality Assessment | Output includes an explicit assessment of user evidence quality using defined criteria: source type, verifiability, recency, and relevance. | Output treats all user evidence as equally valid, or dismisses evidence without quality reasoning. | Parse check: |
Citation Update Accuracy | Corrected response retains valid original citations, removes or marks superseded citations, and adds new citations for user-provided evidence. | Output drops all original citations, keeps a citation that was directly contradicted, or fails to cite user evidence. | Citation diff check: Compare citation list before and after correction. Superseded citations must be flagged, not silently dropped. |
Correction Completeness | All downstream claims that depended on the corrected fact are updated. No orphaned statements remain that assume the original incorrect claim. | Output corrects the specific claim but leaves related statements unchanged, creating internal inconsistency. | LLM-as-judge check: Extract all factual claims from corrected response. Verify none contradict the accepted correction. |
Abstention on Unverifiable Evidence | When user evidence cannot be verified against available sources, output acknowledges the limitation, states what is assumed, and offers a path to resolve. | Output confidently accepts or rejects unverifiable user evidence without qualification. | Schema check: |
Tone and Trust Preservation | Correction response is neutral, professional, and non-defensive. Acknowledges the correction without over-apology or deflection. | Output is defensive, dismissive, excessively apologetic, or shifts blame to sources. | Human review on 5% of cases using a 1-5 Likert scale for tone appropriateness. Threshold: mean score >= 4. |
No Evidence Fabrication | Output introduces no new factual claims that are not grounded in either original sources or user-provided evidence. | Output invents bridging facts, synthesizes unsupported claims, or hallucinates citations to resolve conflicts. | Factual consistency check: Extract all claims. Verify each claim has a source grounding tag. Flag untagged claims for human review. |
Enabling Efficiency, Speed & Accuracy
Intelligent Analysis, Decision & Execution
We build AI systems for teams that need search across company data, workflow automation across tools, or AI features inside products and internal software.
Talk to Us
Search across company data
Give teams answers from docs, tickets, runbooks, and product data with sources and permissions.
Useful when people spend too long searching or get different answers from different systems.

Automate internal workflows
Use AI to route work, draft outputs, trigger actions, and keep approvals and logs in place.
Useful when repetitive work moves across multiple tools and teams.

Add AI to products and internal tools
Build assistants, guided actions, or decision support into the software your team or customers already use.
Useful when AI needs to be part of the product, not a separate tool.
Adapt This Prompt
How to adapt
Start with the base template and a single evidence pair. Use a frontier model with default temperature. Skip structured output enforcement initially—just observe whether the model correctly identifies which evidence prevails and produces a coherent correction.
Simplify the prompt by removing the [EVIDENCE_QUALITY_RUBRIC] section and replacing it with a single instruction: "Prefer more recent, more authoritative sources."
Watch for
- The model agreeing with the user's evidence even when it's weaker than the original source
- Overly verbose explanations that bury the correction
- Failure to cite which specific claim was reversed

About the author
Prasad Kumkar
CEO & MD, Inference Systems
Prasad Kumkar is the CEO & MD of Inference Systems and writes about AI systems architecture, LLM infrastructure, model serving, evaluation, and production deployment. Over 5+ years, he has worked across computer vision models, L5 autonomous vehicle systems, and LLM research, with a focus on taking complex AI ideas into real-world engineering systems.
His work and writing cover AI systems, large language models, AI agents, multimodal systems, autonomous systems, inference optimization, RAG, evaluation, and production AI engineering.
Partnered with leading AI, data, and software stack.
How We Work
Custom AI workflows for your Business
One-fit-all AI don't work for modern businesses. At Inferensys, we aim to understand your business & custom requirements; which we use to define most efficient agentic workflows, the data, and the tools for your business.
01
Review the use case
We understand the task, the users, and where AI can actually help.
Read more02
Pick the right approach
We define what needs search, automation, or product integration.
Read more03
Build the first useful version
We implement the part that proves the value first.
Read more04
Improve from there
We add the checks and visibility needed to keep it useful.
Read moreThe first call is a practical review of your use case and the right next step.
Talk to Us