返回 Skill 列表
extension
分类: 内容与媒体无需 API Key

Override Mechanisms

Override Mechanisms允许人类纠正或逆转AI的决策

person作者: jakexiaohubgithub

Override Mechanisms

Skill Profile

(Select at least one profile to enable specific modules)

  • [ ] DevOps
  • [x] Backend
  • [ ] Frontend
  • [ ] AI-RAG
  • [ ] Security Critical

Overview

Override Mechanisms allow humans to correct or reverse AI decisions, providing a critical safety net for automated systems. Proper override implementation includes tracking, justification, learning, and prevention of abuse.

Core Principle: "AI should be overridable, but overrides should be logged, justified, and learned from."


Why This Matters

  • <Benefit>: <short explanation>
  • <Benefit>: <short explanation>
  • <Benefit>: <short explanation>

Core Concepts & Rules

1. Core Principles

  • Follow established patterns and conventions
  • Maintain consistency across codebase
  • Document decisions and trade-offs

2. Implementation Guidelines

  • Start with the simplest viable solution
  • Iterate based on feedback and requirements
  • Test thoroughly before deployment

Inputs / Outputs / Contracts

  • Inputs:
    • <e.g., env vars, request payload, file paths, schema>
  • Entry Conditions:
    • <Pre-requisites: e.g., Repo initialized, DB running, specific branch checked out>
  • Outputs:
    • <e.g., artifacts (PR diff, docs, tests, dashboard JSON)>
  • Artifacts Required (Deliverables):
    • <e.g., Code Diff, Unit Tests, Migration Script, API Docs>
  • Acceptance Evidence:
    • <e.g., Test Report (screenshot/log), Benchmark Result, Security Scan Report>
  • Success Criteria:
    • <e.g., p95 < 300ms, coverage ≥ 80%>

Skill Composition

  • Depends on: None
  • Compatible with: None
  • Conflicts with: None
  • Related Skills: None

Quick Start

Assumptions

  • Users have appropriate permissions for their role
  • Override reasons are provided in good faith
  • Override correctness can be verified
  • Model can be improved from override data

Compatibility

  • Works with any AI/ML system
  • Language-agnostic override patterns
  • Can integrate with existing permission systems

Test Scenario Matrix

| Scenario | Override Type | Expected Behavior | Notes | |----------|--------------|-------------------|-------| | Low-impact decision | Manual | Immediate override | No approval needed | | High-impact decision | Manual | Requires approval | Manager must approve | | VIP customer | Business rule | Auto-approve | Rule-based override | | Emergency | Emergency | Kill switch | Immediate, alerts team | | Suspicious pattern | Abuse detection | Alert manager | Prevent bulk overrides |


Technical Guardrails & Security Threat Model

1. Security & Privacy (Threat Model)

  • Top Threats: Injection attacks, authentication bypass, data exposure
  • [ ] Data Handling: Sanitize all user inputs to prevent Injection attacks. Never log raw PII
  • [ ] Secrets Management: No hardcoded API keys. Use Env Vars/Secrets Manager
  • [ ] Authorization: Validate user permissions before state changes

2. Performance & Resources

  • [ ] Execution Efficiency: Consider time complexity for algorithms
  • [ ] Memory Management: Use streams/pagination for large data
  • [ ] Resource Cleanup: Close DB connections/file handlers in finally blocks

3. Architecture & Scalability

  • [ ] Design Pattern: Follow SOLID principles, use Dependency Injection
  • [ ] Modularity: Decouple logic from UI/Frameworks

4. Observability & Reliability

  • [ ] Logging Standards: Structured JSON, include trace IDs request_id
  • [ ] Metrics: Track error_rate, latency, queue_depth
  • [ ] Error Handling: Standardized error codes, no bare except
  • [ ] Observability Artifacts:
    • Log Fields: timestamp, level, message, request_id
    • Metrics: request_count, error_count, response_time
    • Dashboards/Alerts: High Error Rate > 5%

Agent Directives & Error Recovery

(ข้อกำหนดสำหรับ AI Agent ในการคิดและแก้ปัญหาเมื่อเกิดข้อผิดพลาด)

  • Thinking Process: Analyze root cause before fixing. Do not brute-force.
  • Fallback Strategy: Stop after 3 failed test attempts. Output root cause and ask for human intervention/clarification.
  • Self-Review: Check against Guardrails & Anti-patterns before finalizing.
  • Output Constraints: Output ONLY the modified code block. Do not explain unless asked.

Definition of Done

  • [ ] Override permissions defined and implemented
  • [ ] Justification requirements enforced
  • [ ] Approval workflows for high-impact decisions
  • [ ] Comprehensive logging of all overrides
  • [ ] Feedback loop to model training
  • [ ] Abuse detection and alerting
  • [ ] Emergency override mechanism
  • [ ] Override analytics dashboard
  • [ ] Integration tests passing
  • [ ] Documentation complete

Anti-patterns / Pitfalls

  • Don't: Log PII, catch-all exception, N+1 queries
  • ⚠️ Watch out for: Common symptoms and quick fixes
  • 💡 Instead: Use proper error handling, pagination, and logging

Reference Links


Versioning & Changelog

  • Version: 1.0.0
  • Changelog:
    • 2026-02-22: Initial version with complete template structure