Inferensys

Prompt

Quote Fabrication Risk Assessment Prompt

A practical prompt playbook for using the Quote Fabrication Risk Assessment Prompt in production AI workflows. Designed for investigative and editorial teams who need to triage suspicious quotes before human review.
Risk analyst performing AI risk assessment on laptop, risk matrices visible, casual office risk session.
PROMPT PLAYBOOK

When to Use This Prompt

Defines the job-to-be-done, ideal user, required context, and boundaries for the Quote Fabrication Risk Assessment Prompt.

This prompt is a triage instrument for editorial and investigative teams who need to evaluate the likelihood that an attributed quote is fabricated. It is not a final verdict. The ideal user is a researcher, journalist, or fact-checker who has encountered a quote in a draft article, social media post, or third-party report and needs a structured, repeatable way to assess its authenticity before investing hours in manual investigation. The prompt requires three pieces of context to function: the quote itself, information about the attributed speaker (such as their known speech patterns, public role, and typical venues), and any available provenance data (such as the publication where the quote appeared, a claimed date, or an archival reference). Without these inputs, the model cannot produce a calibrated risk score.

The prompt analyzes the quote against the provided context to produce a risk score and supporting indicators. These indicators include stylistic mismatch (does the quote sound like the attributed speaker?), absence from known archives (has the quote appeared in any verifiable record?), and provenance gaps (is the chain of custody for the quote broken or suspicious?). The output is designed to be consumed by a human decision-maker or routed into a downstream verification pipeline. High-risk items should always route to a human investigator. This prompt does not perform source-to-text comparison, which is handled by sibling prompts in the Quote Accuracy and Paraphrase Fidelity Verification group, such as the Quote-to-Source Comparison Prompt Template. It also does not search the web or query external databases; it reasons only over the context you provide.

Do not use this prompt when you need a definitive yes/no answer on fabrication. It is a risk assessment tool, not a lie detector. Do not use it when the attributed speaker is unknown or the provenance information is entirely absent—the model will have no signal to work with and may produce a misleadingly confident score. Do not use it as a replacement for source-to-text comparison when the original source is available; use the Quote-to-Source Comparison Prompt instead. For production systems, always pair this prompt with a human review step for any output above a 'medium' risk threshold, and log every assessment with the input context, model version, and risk score for auditability.

PRACTICAL GUARDRAILS

Use Case Fit

Where the Quote Fabrication Risk Assessment Prompt works, where it fails, and what inputs it assumes.

01

Good Fit: Triage for High-Volume Editorial

Use when: you have a large queue of attributed quotes and need to surface the riskiest items for human investigators. Guardrail: set a low risk-score threshold for human review to avoid false negatives; let the prompt be sensitive rather than specific.

02

Bad Fit: Definitive Fabrication Verdicts

Avoid when: the output will be used as a final determination without human review. Guardrail: the prompt produces a risk score, not a verdict. Always route high-risk items to a human before publication or action.

03

Required Inputs

What you must provide: the attributed quote text, the claimed speaker, the alleged date and context of the statement, and any available source archives or retrieval tool access. Guardrail: missing provenance fields should force a higher risk score, not a guess.

04

Operational Risk: Stylometric Overconfidence

What to watch: the model may over-weight stylistic mismatch as fabrication evidence when the speaker's style varies by context, audience, or time period. Guardrail: weight stylistic indicators lower than provenance gaps and archive absence. Include a calibration note in the output.

05

Operational Risk: Temporal Blind Spots

What to watch: the model cannot verify quotes against sources outside its training data or retrieval window. Guardrail: flag any quote where the alleged date falls after the model's knowledge cutoff or retrieval index freshness. Escalate to human with a 'temporal gap' indicator.

06

Integration: Downstream Routing

What to watch: risk scores without routing logic create bottlenecks. Guardrail: wire the output risk score and indicator list into a triage system that auto-escalates high-risk items, queues medium-risk for batch review, and logs low-risk for audit sampling.

PROMPT PLAYBOOK

Copy-Ready Prompt Template

A copy-ready prompt for triaging attributed quotes by fabrication risk, with placeholders for the quote, source context, and risk thresholds.

The prompt below is designed to be pasted directly into your system prompt or user message field. It instructs the model to act as a fabrication-risk triage analyst, evaluating a single attributed quote against a set of indicators and producing a structured risk score. The output is intended to route high-risk items to human investigators, not to make a final determination of fabrication. Before using this prompt, ensure you have the quote text, the claimed speaker, the publication or context where it appeared, and any available provenance metadata.

text
You are a quote-fabrication risk analyst. Your job is to triage attributed quotes by estimating the likelihood that they were fabricated, partially fabricated, or misattributed. You do not make final determinations. You produce a structured risk assessment that routes high-risk items to human investigators.

## INPUT
- Quote: [QUOTE_TEXT]
- Claimed Speaker: [SPEAKER_NAME]
- Source Context: [SOURCE_DESCRIPTION]
- Publication Date: [PUBLICATION_DATE]
- Available Provenance: [PROVENANCE_NOTES]

## OUTPUT SCHEMA
Return a single JSON object with the following fields:
- "fabrication_risk_score": number between 0.0 and 1.0, where 0.0 is almost certainly authentic and 1.0 is almost certainly fabricated.
- "risk_category": one of "LOW", "MEDIUM", "HIGH", "CRITICAL".
- "primary_indicators": array of strings naming the strongest indicators that raised the risk score.
- "indicator_details": array of objects, each with "indicator" (string), "finding" (string), and "weight" (number 0.0-1.0).
- "recommended_action": one of "CLEAR", "REVIEW", "ESCALATE", "QUARANTINE".
- "investigator_brief": string summarizing what a human investigator should examine first.
- "abstention": boolean. Set to true if you cannot produce a reliable assessment (e.g., insufficient information).
- "abstention_reason": string explaining why you abstained, if applicable.

## INDICATORS TO EVALUATE
1. **Provenance Absence**: No record of the quote in archives, transcripts, or prior publications.
2. **Stylistic Mismatch**: Vocabulary, syntax, or phrasing inconsistent with the claimed speaker's known style.
3. **Temporal Implausibility**: Quote references events, terminology, or knowledge unavailable at the claimed time.
4. **Source Chain Break**: The quote appears only in a single unverifiable source with no upstream attribution.
5. **Convenience Bias**: The quote aligns too perfectly with the quoter's argument while contradicting the speaker's known positions.
6. **Anomalous Specificity**: Unusual precision or detail that is absent from the speaker's verified statements.
7. **Duplicate or Recycled Content**: Quote matches or closely mirrors a known quote from a different speaker or context.

## CONSTRAINTS
- Do not assume fabrication from a single weak indicator. Weigh multiple indicators together.
- If provenance information is missing, note it as an indicator rather than abstaining immediately.
- Distinguish between "likely fabricated" and "unverifiable." Use the abstention field when evidence is too thin to score.
- Do not output markdown. Return only the JSON object.

To adapt this prompt, replace each square-bracket placeholder with real values before sending. The [QUOTE_TEXT] should be the exact text as published. [SPEAKER_NAME] is the person to whom the quote is attributed. [SOURCE_DESCRIPTION] should include the publication name, article title, and any surrounding context. [PUBLICATION_DATE] helps the model assess temporal plausibility. [PROVENANCE_NOTES] should capture what you already know about the quote's origin—such as whether it appears in transcripts, archives, or prior reporting—or state 'No provenance information available.' The indicator list can be extended or pruned based on your domain. For high-stakes domains such as legal or political fact-checking, add domain-specific indicators such as 'Legal Terminology Misuse' or 'Policy Position Contradiction.' The output schema is designed for direct ingestion by a triage router, so keep field names stable across adaptations.

IMPLEMENTATION TABLE

Prompt Variables

Inputs the Quote Fabrication Risk Assessment Prompt needs to work reliably. Validation notes describe what makes each variable usable.

PlaceholderPurposeExampleValidation Notes

[QUOTE_TEXT]

The attributed quote under investigation

"We have seen a 300% increase in productivity since implementing the new system."

Must be a non-empty string. Check for minimum length (e.g., 20 characters) to avoid null inputs. Trim whitespace.

[ATTRIBUTED_SPEAKER]

The person or entity the quote is attributed to

Jane Doe, CEO of Acme Corp

Must be a non-empty string. Validate against a known entity list if available. Flag generic titles without names (e.g., 'a company spokesperson') as higher risk.

[SOURCE_DOCUMENT]

The full text of the document containing the quote

Full text of a press release, article, or social media post.

Must be a non-empty string. Check for minimum length. If the quote is not found verbatim within [SOURCE_DOCUMENT], log a pre-processing error before assessment.

[PUBLICATION_DATE]

The date the source document was published or accessed

2024-03-15

Must be a valid ISO 8601 date string (YYYY-MM-DD). Null allowed if genuinely unknown, but this should increase the risk score for 'temporal inconsistency'.

[CORROBORATION_SOURCES]

A list of other sources where the quote or similar statements appear

Must be a valid JSON array of URL strings. An empty array is a strong fabrication risk indicator. Validate URLs are well-formed.

[SPEAKER_STYLE_REFERENCE]

A sample of verified text written or spoken by the attributed speaker

"Our quarterly results show steady, incremental progress."

Must be a non-empty string of at least 100 characters for a meaningful stylistic comparison. Null allowed, but disables the stylistic mismatch check and should be noted in the output.

[CONTEXT_WINDOW]

Surrounding text from the [SOURCE_DOCUMENT] for context-stripping analysis

The two paragraphs immediately before and after the quote in the source document.

Must be a string. Null allowed if the quote is the entire document. If provided, its length should be significantly larger than [QUOTE_TEXT] to be useful.

PROMPT PLAYBOOK

Implementation Harness Notes

How to wire the Quote Fabrication Risk Assessment Prompt into a production triage pipeline with validation, tool integration, and human review routing.

The Quote Fabrication Risk Assessment Prompt is designed as a triage classifier, not a final arbiter. Its job is to score the likelihood that an attributed quote was fabricated and route high-risk items to human investigators. In a production harness, this prompt sits between quote extraction and human review. It should never auto-publish a fabrication verdict. The harness must enforce that outputs above a configurable risk threshold are always escalated, and even low-risk scores should be logged for audit and trend analysis.

Wire the prompt into an application with a strict JSON output contract and a post-processing validation layer. The model must return a structured object containing at minimum: fabrication_risk_score (0.0–1.0), risk_indicators (a list of named indicators with confidence and evidence excerpts), provenance_status (e.g., verified_source, unverifiable, no_record), and a recommendation field constrained to auto_pass, review, or escalate. Use a JSON schema validator in your application layer—not just the prompt—to reject malformed responses and trigger a retry with the error message appended to the next request. For high-stakes editorial workflows, add a human-in-the-loop gate: any output with fabrication_risk_score >= 0.6 or recommendation == 'escalate' must be queued for review with the full prompt input, model output, and retrieved evidence packaged into a review packet.

The prompt works best when it has access to external evidence. Integrate it with a search or archive retrieval tool (e.g., a news database API, web search, or internal transcript store) before calling the model. The harness should: (1) extract the quote and attribution from the source content, (2) query external sources for the quote string and speaker-context pairs, (3) inject retrieved snippets into the [EVIDENCE] placeholder, and (4) call the model. If no evidence is found, the prompt's provenance_status will naturally flag no_record, which is a strong fabrication indicator. Log every step—extraction, retrieval queries, retrieved documents, model input, model output, and final routing decision—for downstream audit and prompt regression testing. Avoid calling the model without evidence; the prompt's value degrades significantly when it can only reason about stylistic mismatch without source-grounding.

For model choice, prefer models with strong reasoning and structured output capabilities. The prompt requires multi-factor analysis (provenance, style, temporal consistency, archive presence) and calibrated uncertainty, which smaller or older models may handle inconsistently. If using a model that doesn't support strict JSON mode natively, implement a retry loop with format-correction prompts from the Output Repair and Validation pillar. Set a maximum of 2 retries before logging the failure and routing to human review with a model_failure flag. For batch processing across many quotes, implement rate limiting, cost tracking per quote, and a dead-letter queue for items that exhaust retries.

Failure modes to monitor in production: (1) the model overweights stylistic mismatch for non-native speakers or translated quotes, producing false-positive fabrication flags; (2) the model underweights fabrication risk when a quote appears in one low-authority source that was indexed but not verified; (3) provenance checks fail silently when the retrieval system returns empty results due to query formulation errors rather than genuine absence. Mitigate these by running a weekly calibration set of known-fabricated and known-genuine quotes through the harness and tracking precision/recall at your chosen threshold. If the fabrication risk score distribution drifts (e.g., sudden spike in high-risk flags), investigate retrieval pipeline health before adjusting the prompt.

Next steps after implementation: build a review dashboard that shows investigators the quote, the model's risk score and indicators, the retrieved evidence, and a quick accept/override interface. Log every human decision to create a feedback dataset for future fine-tuning or prompt improvement. Do not use the prompt's output as a training label without human verification—the prompt is a triage tool, not ground truth.

IMPLEMENTATION TABLE

Expected Output Contract

Fields, data types, and validation rules for the JSON response returned by the Quote Fabrication Risk Assessment Prompt. Use this contract to build a parser, validator, and retry logic in your application harness.

Field or ElementType or FormatRequiredValidation Rule

risk_score

number (0.0 to 1.0)

Must be a float between 0 and 1 inclusive. Parse check: JSON number type. Threshold check: values above 0.8 should trigger human review routing.

risk_level

string (enum)

Must be one of: 'low', 'medium', 'high', 'critical'. Schema check: exact string match. Must be consistent with risk_score thresholds defined in the prompt instructions.

fabrication_indicators

array of objects

Each object must contain 'indicator' (string), 'severity' (enum: 'minor', 'moderate', 'strong'), and 'evidence' (string). Array must not be empty if risk_level is 'high' or 'critical'.

provenance_status

string (enum)

Must be one of: 'verified', 'unverifiable', 'contradicted', 'no_record'. Schema check: exact string match. 'no_record' should force risk_level to at least 'high'.

stylometric_anomaly_score

number (0.0 to 1.0) or null

If present, must be a float between 0 and 1. Null allowed when quote is too short for analysis. Parse check: JSON number or null. Values above 0.7 should be flagged in fabrication_indicators.

source_attestation_count

integer

Must be a non-negative integer. Parse check: JSON number without decimal. Zero attestations must co-occur with provenance_status 'no_record' or 'unverifiable'.

recommended_action

string (enum)

Must be one of: 'publish', 'review', 'escalate', 'retract'. Schema check: exact string match. 'escalate' required when risk_level is 'critical'. 'publish' only allowed when risk_level is 'low'.

investigation_priority

string (enum)

Must be one of: 'routine', 'elevated', 'urgent'. Schema check: exact string match. Must be 'urgent' when risk_score is above 0.9 or provenance_status is 'contradicted'.

PRACTICAL GUARDRAILS

Common Failure Modes

Quote fabrication risk assessment is a high-stakes triage task. These are the most common failure modes observed when deploying this prompt in production editorial and investigative workflows, along with concrete mitigations.

01

Stylometric Overfitting to a Single Corpus

What to watch: The model learns a narrow stylistic fingerprint from a small sample of the speaker's known quotes and flags any deviation as high-risk, even when the speaker's style naturally varies by audience, medium, or time period. Guardrail: Provide a diverse calibration set spanning multiple contexts, formats, and time periods. Include explicit instructions to weigh stylistic mismatch as one indicator among many, not a dispositive signal.

02

Provenance Gap Misclassification

What to watch: The model conflates 'no source found in available archives' with 'likely fabricated,' generating high-risk scores for quotes that are simply obscure, recently published, or behind paywalls. Guardrail: Require the prompt to distinguish between 'absence of evidence' and 'evidence of absence.' Output a separate provenance-gap flag and route gap-only cases to human search escalation rather than auto-flagging as fabrication.

03

Temporal Context Blindness

What to watch: The model fails to account for when a quote was allegedly made, comparing it against anachronistic sources or flagging terminology that was common at the time but is now outdated. Guardrail: Include the alleged date or date range as a required input field. Instruct the model to restrict archival checks and stylistic comparisons to the relevant time window.

04

Confidence Inflation on Ambiguous Cases

What to watch: The model assigns a high-confidence 'likely fabricated' or 'likely authentic' score to quotes with thin or contradictory evidence, failing to express appropriate uncertainty. Guardrail: Add a calibration instruction requiring the model to reduce confidence when evidence is sparse, conflicting, or circumstantial. Implement a post-processing threshold that routes medium-confidence outputs to human review regardless of the predicted label.

05

Adversarial Quote Injection Bypass

What to watch: A fabricated quote is embedded in a document that otherwise contains verifiable, authentic quotes from the same speaker, causing the model to anchor on the surrounding legitimacy and miss the injection. Guardrail: Instruct the model to evaluate each quote independently, not relative to neighboring quotes. Include few-shot examples of injection patterns and require per-quote provenance checks even when the document appears authoritative.

06

Translation and Dialect Mismatch Flagging

What to watch: The model flags a quote as stylistically mismatched because it was translated from another language or reflects a regional dialect variant, misinterpreting translation artifacts as fabrication signals. Guardrail: Add a metadata field for original language and translation status. Instruct the model to suppress stylistic mismatch flags when translation is confirmed and to note translation-related uncertainty separately.

IMPLEMENTATION TABLE

Evaluation Rubric

Run these checks against a golden dataset of quotes with known fabrication status. Each criterion targets a specific failure mode of the Quote Fabrication Risk Assessment Prompt.

CriterionPass StandardFailure SignalTest Method

Fabrication Detection Recall

Risk score >= 0.8 for all known fabricated quotes in the golden set

Low risk score on a known fabricated quote; fabricated quote classified as low-risk

Run prompt on 50+ confirmed fabricated quotes; measure recall at threshold 0.8

False-Positive Rate on Verified Quotes

Risk score <= 0.3 for all verified authentic quotes with strong provenance

High risk score assigned to a quote with documented source, date, and recording

Run prompt on 50+ verified authentic quotes; measure false-positive rate at threshold 0.3

Provenance Gap Flagging

Provenance gap indicator is true when no source, date, or recording exists in [INPUT]

Provenance gap indicator is false despite missing all three provenance fields

Provide quotes with systematically removed provenance fields; check indicator output

Stylistic Mismatch Detection

Stylistic mismatch indicator is true when quote vocabulary and syntax differ significantly from [SPEAKER_STYLE_REFERENCE]

Indicator is false despite clear lexical and syntactic divergence from known speaker patterns

Pair fabricated quotes with authentic speaker style samples; verify mismatch flag triggers

Archive Absence Handling

Archive absence indicator is true when quote is not found in [ARCHIVE_SOURCE] and confidence is appropriately lowered

Confidence remains high despite confirmed absence from searchable archives

Use quotes known to be absent from specified archives; verify indicator and confidence drop

Confidence Calibration

Output confidence score correlates with actual fabrication likelihood across the golden set

Confidence scores are uniformly high or show no correlation with ground-truth labels

Plot confidence scores against binary fabrication labels; compute Expected Calibration Error

Abstention on Insufficient Evidence

Prompt returns abstention=true or risk_score=null when [INPUT] contains only the quote with no context

Prompt assigns a high-confidence score with no provenance, style, or archive data to evaluate

Provide bare quote strings with zero metadata; verify abstention or null risk score

Output Schema Compliance

Output matches [OUTPUT_SCHEMA] exactly; all required fields present with correct types

Missing risk_score, indicators, or rationale fields; type mismatch in confidence field

Validate output against JSON Schema; run on 100 varied inputs; require 100% schema compliance

ADAPTATION OPTIONS

Adapt This Prompt

How to adapt

Use the base prompt with a frontier model (GPT-4o, Claude 3.5 Sonnet) and manual review of every output. Skip structured output enforcement initially—let the model return markdown with risk scores and indicators. Focus on calibrating the risk dimensions (provenance, stylistic match, archival presence) against a small hand-labeled dataset of known fabricated and genuine quotes.

Watch for

  • Overconfident scores without hedging language
  • Missing indicator categories (e.g., flagging stylistic mismatch but not provenance gaps)
  • Model treating absence of evidence as evidence of fabrication
Prasad Kumkar

About the author

Prasad Kumkar

CEO & MD, Inference Systems

Prasad Kumkar is the CEO & MD of Inference Systems and writes about AI systems architecture, LLM infrastructure, model serving, evaluation, and production deployment. Over 5+ years, he has worked across computer vision models, L5 autonomous vehicle systems, and LLM research, with a focus on taking complex AI ideas into real-world engineering systems.

His work and writing cover AI systems, large language models, AI agents, multimodal systems, autonomous systems, inference optimization, RAG, evaluation, and production AI engineering.