# AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

Source: [arXiv](https://arxiv.org/abs/2607.28617v1)  
Feed7 permalink: https://feed7.dev/p/2607-28617v1-1cbh1qu  
Published: 2026-07-30T17:58:58.000Z  
Trust: Needs Review (needs_review)

## Why Included

AISPA turns system-prompt review into an eight-dimension audit. Its survey suggests builders should test prompts for user protection and conflicting instructions, not merely check that safeguards exist.

## Source Summary

AISPA classifies **3,249 instructions** from **88 commercial AI products** as protective or problematic across eight user-centered dimensions. Although 98.9% of products include a protection, only 24% cover every dimension.

## Practical Implication

Audit agent system prompts instruction by instruction, checking both coverage and conflicts. A long safety section is weak evidence when protective and user-hostile directives can coexist in the same prompt.

## Agent-Ready Context

AISPA classifies **3,249 instructions** from **88 commercial AI products** as protective or problematic across eight user-centered dimensions. Although 98.9% of products include a protection, only 24% cover every dimension.

Audit agent system prompts instruction by instruction, checking both coverage and conflicts. A long safety section is weak evidence when protective and user-hostile directives can coexist in the same prompt.

The study reports that roughly **40% of products** contain at least one problematic instruction, but the abstract does not establish how its taxonomy transfers to private coding-agent harnesses or predicts runtime behavior.

## Connected Context

Feed7 judgment across 297 accumulated Signals:

AISPA turns prompt review into a user-centered, instruction-level audit: broad safety language is insufficient when protections omit dimensions or coexist with harmful directives. Against the prior candidates, it adds a static governance check that complements—but cannot replace—task-level evaluation of runtime behavior and context sensitivity.

- [asgeirtj/system_prompts_leaks](https://feed7.dev/p/system-prompts-leaks-0pth2c5) — The prompt archive offers real product prompts to which AISPA’s instruction-level taxonomy could be applied, while AISPA supplies a structured audit method beyond informal comparison.
- [How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads](https://feed7.dev/p/how-evals-and-prompts-shape-agent-behavior-preetika-bhateja-daniel-bump-1cmecaw) — AISPA’s static conflict and coverage audit complements the production loop of trace review and evals, which is still needed because prompt contents alone do not establish runtime behavior.
- [The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context](https://feed7.dev/p/2607-12963v1-1oc0qmr) — The finding that irrelevant context can flip individual outputs reinforces evaluating audited prompts per task rather than assuming improved instruction coverage guarantees stable behavior.

## Context Map

- Layer: context
- Domains: coding
- Topics: prompting, context-engineering, agent-reliability

## Uncertainty

- The study reports that roughly **40% of products** contain at least one problematic instruction, but the abstract does not establish how its taxonomy transfers to private coding-agent harnesses or predicts runtime behavior.

## Agent Instruction

Use this item as source-backed context. Do not invent claims beyond the linked source. If this item conflicts with another source, call out the conflict.
