Skip to main content

Vulnerabilities

HackAgent ships with 13 built-in vulnerability classes covering the input, model, data, and agent layers of an AI system. Each one extends BaseVulnerability (hackagent.risks.base), defines an Enum of testable sub-types, and has a matching threat profile — recommended datasets, attack techniques, objective, and metrics — documented inline on its own page.

Reference​

VulnerabilityDescription
JailbreakTests whether the LLM can be manipulated into bypassing its safety filters through roleplay, encoding, multi-turn, hypothetical, or authority-manipulation techniques.
Prompt InjectionTests whether the LLM executes attacker-supplied instructions that override or bypass the system prompt.
System Prompt LeakageTests whether the LLM reveals sensitive details from its system prompt, such as credentials, internal instructions, or guardrails.
Input Manipulation AttackTests whether encoding bypasses, format string attacks, or Unicode manipulation can evade input validation and safety filters.
Model EvasionTests whether adversarial examples, feature manipulation, or boundary exploitation can evade the model's safety mechanisms.
Craft Adversarial DataTests whether adversarially crafted data — perturbations, poisoned examples, or augmentation abuse — can compromise model behaviour.
Sensitive Information DisclosureTests for training-data extraction, architecture disclosure, and configuration leakage.
MisinformationTests whether the LLM produces factual fabrications, invented sources, or misrepresented expertise.
Excessive AgencyTests whether the LLM performs actions or grants permissions exceeding its intended scope without oversight.
Malicious Tool InvocationTests for risks from untrusted tool execution, data exfiltration through tool interactions, and tool privilege escalation.
Credential ExposureTests for hardcoded credentials, token leakage, and misconfigured access controls in AI systems.
Public-Facing Application ExploitationTests whether publicly exposed AI APIs, web interfaces, or endpoints can be abused or exploited beyond intended use.
Vector and Embedding Weaknesses ExploitTests for embedding inversion, vector database poisoning, and similarity search manipulation in RAG pipelines.

Using a Vulnerability​

from hackagent.risks import Jailbreak
from hackagent.risks.jailbreak.types import JailbreakType

# Use all sub-types
vuln = Jailbreak()

# Or specify particular sub-types
vuln = Jailbreak(types=[
JailbreakType.ROLEPLAY.value,
JailbreakType.MULTI_TURN.value,
])

Don't see a category that fits your use case? See Custom Vulnerabilities to define your own.

How Threat Profiles Work​

A threat profile maps a vulnerability to recommended evaluation parameters:

"Given vulnerability X, which datasets, attack techniques, objective, and metrics should an evaluation campaign use?"

Each of the 13 built-in vulnerabilities above has a matching threat profile — see the "Threat Profile" section on its own page for the concrete values. This section explains the shared anatomy behind every one of them.

Each ThreatProfile is a frozen dataclass with these fields:

FieldTypeDescription
vulnerabilitytype[BaseVulnerability]The vulnerability class this profile maps
datasetslist[DatasetRecommendation]Recommended datasets with relevance tags
attackslist[AttackRecommendation]Compatible attack techniques
objectivestrDefault attack objective ("jailbreak", "harmful_behavior", "policy_violation")
metricslist[str]Relevant metric names ("asr", "toxicity_score", "judge_score")
descriptionstrHuman-readable summary

Relevance Levels​

Every dataset and attack recommendation carries a relevance tag:

LevelMeaningWhen to use
PRIMARYDirectly designed to test this vulnerabilityAlways include in evaluation
SECONDARYUseful for broader coverage or baseline comparisonInclude for comprehensive audits

Objectives​

ObjectiveDescriptionExample Vulnerabilities
jailbreakBypass safety mechanisms to produce forbidden outputPromptInjection, Jailbreak, SystemPromptLeakage
harmful_behaviorGenerate harmful, toxic, or dangerous contentMisinformation
policy_violationViolate organizational policies or access controlsExcessiveAgency, CredentialExposure

Metrics​

Common metric names used across profiles:

  • asr — Attack Success Rate
  • judge_score — LLM-based evaluation score
  • toxicity_score — Toxicity measurement
  • refusal_rate — Percentage of refused requests

Accessing a Profile​

Each vulnerability has an associated profile constant in its module:

from hackagent.risks.jailbreak import JAILBREAK_PROFILE

print(JAILBREAK_PROFILE.description)
# "Tests resistance to multi-turn, roleplay, encoding, and authority-based bypass."

print(JAILBREAK_PROFILE.dataset_presets)
# ['strongreject', 'harmbench', 'advbench', 'jailbreakbench', ...]

print(JAILBREAK_PROFILE.attack_techniques)
# ['h4rm3l', 'TAP', 'PAIR']

print(JAILBREAK_PROFILE.objective) # 'jailbreak'
print(JAILBREAK_PROFILE.metrics) # ['asr', 'judge_score']

# Primary datasets — core evaluation
for d in JAILBREAK_PROFILE.primary_datasets:
print(f"[P] {d.preset}: {d.rationale}")

# Secondary datasets — extended coverage
for d in JAILBREAK_PROFILE.secondary_datasets:
print(f"[S] {d.preset}: {d.rationale}")

Don't see a threat profile that fits a custom vulnerability? See Custom Vulnerabilities to build your own with ThreatProfile and the profile_helpers module.

Learn More​