hackagent.agent
HackAgent Objects
class HackAgent()
The primary client for orchestrating security assessments with HackAgent.
This class serves as the main entry point to the HackAgent library, providing a high-level interface for:
- Configuring victim agents that will be assessed.
- Defining and selecting attack strategies.
- Executing automated security tests against the configured agents.
- Retrieving and handling test results.
It encapsulates complexities such as agent registration
with the local backend (via AgentRouter), and the dynamic dispatch of various
attack methodologies.
Attributes:
router- AnAgentRouterinstance managing the agent's representation in the HackAgent backend.attack_strategies- A dictionary mapping strategy names to theirAttackStrategyimplementations.
__init__
def __init__(endpoint: str,
name: Optional[str] = None,
agent_type: Union[AgentTypeEnum, str] = AgentTypeEnum.UNKNOWN,
base_url: Optional[str] = None,
api_key: Optional[str] = None,
raise_on_unexpected_status: bool = False,
timeout: Optional[float] = 120.0,
metadata: Optional[Dict[str, Any]] = None,
target_config: Optional[Dict[str, Any]] = None,
adapter_operational_config: Optional[Dict[str, Any]] = None,
thinking: Optional[bool] = None,
before_guardrail: Optional[Dict[str, Any]] = None,
after_guardrail: Optional[Dict[str, Any]] = None)
Initializes the HackAgent client and prepares it for interaction.
This constructor sets up the local storage backend, loads default prompts, resolves the agent type, and initializes the agent router to ensure the agent is known to the backend. It also prepares available attack strategies.
Arguments:
endpoint- The target application's endpoint URL. This is the primary interface that the configured agent will interact with or represent during security tests.name- An optional descriptive name for the agent being configured. If not provided, a default name might be assigned or behavior might depend on the specific backend agent management policies.agent_type- Specifies the type of the agent. This can be provided as anAgentTypeEnummember (e.g.,AgentTypeEnum.GOOGLE_ADK) or as a string identifier (e.g., "google-adk", "litellm"). String values are automatically converted to the correspondingAgentTypeEnummember. Defaults toAgentTypeEnum.UNKNOWNif not specified or if an invalid string is provided.raise_on_unexpected_status- If set toTrue, the API client will raise an exception for any HTTP status codes that are not typically expected for a successful operation. Defaults toFalse.timeout- The timeout duration in seconds for API requests made by the authenticated (remote) HackAgent backend client. Defaults to120.0seconds so requests to a misbehaving/unreachable backend fail predictably instead of hanging indefinitely. PassNoneexplicitly to opt out and disable the timeout (unbounded wait, the previous default behavior).metadata- Optional dictionary containing agent-specific metadata.target_config- Optional default request settings for the configured victim model. This is the preferred place to define target-side generation defaults such asmax_tokens,temperature, andtimeout.adapter_operational_config- Optional configuration for the agent adapter.thinking- Optional OLLAMA-only control for reasoning traces. When set toFalse, requests sent through the target OLLAMA adapter includethink: falseto disable thinking output. Ignored for non-OLLAMA target agent types.
attack_strategies
@property
def attack_strategies() -> Dict[str, Any]
Lazy-loaded attack strategies dictionary.
hack
def hack(attack_config: Dict[str, Any],
run_config_override: Optional[Dict[str, Any]] = None,
fail_on_run_error: bool = True,
_tui_event_bus: Optional[Any] = None) -> Any
Executes a specified attack strategy against the configured victim agent.
This method serves as the primary action command for initiating an attack.
It identifies the appropriate attack strategy based on attack_config,
ensures the victim agent (managed by self.router) is ready, and then
delegates the execution to the chosen strategy.
Arguments:
attack_config- A dictionary containing parameters specific to the chosen attack type. Must include an 'attack_type' key that maps to a registered strategy (e.g., "advprefix"). Other keys provide configuration for that strategy (e.g., 'category', 'prompt_text').run_config_override- An optional dictionary that can override default run configurations. The specifics depend on the attack strategy and backend capabilities.fail_on_run_error- IfTrue(the default), an exception will be raised if the attack run encounters an error and fails. IfFalse, errors might be suppressed or handled differently by the strategy.
Returns:
The result returned by the execute method of the chosen attack
strategy. The nature of this result is strategy-dependent.
Raises:
ValueError- If the 'attack_type' is missing fromattack_configor if the specified 'attack_type' is not a supported/registered strategy.HackAgentError- For issues during backend agent operations, or other unexpected errors during the attack process.
hack_chain
def hack_chain(attacks: Optional[list] = None,
goals: Optional[list] = None,
run_config_override: Optional[Dict[str, Any]] = None,
fail_on_run_error: bool = True,
escalate_only_mitigated: bool = True,
_tui_event_bus: Optional[Any] = None) -> list
Runs a sequence of attack strategies against a shared pool of goals.
By default (escalate_only_mitigated=True) this implements a
"fallback ladder": every goal starts at attacks[0]. Any goal for
which the victim's response is judged successful (a jailbreak/
violation) is considered resolved and is dropped from the chain — it
is never retried. Any goal that is mitigated (the victim's response
is judged safe) is carried over and retried with attacks[1], then
attacks[2], and so on, until either the goal succeeds or the
chain is exhausted.
With escalate_only_mitigated=False, every goal is instead sent to
every attack in the chain regardless of outcome — useful for
running several attacks against the same goal set and collecting all
of their results in one call, rather than escalating only failures.
Success/mitigation is determined per goal from the evaluated result
rows returned by each step (see
hackagent.attacks.evaluator.metrics.is_successful_result): a goal
is considered successful for a step if any of its result rows for
that step are judged successful.
Arguments:
attacks- Ordered list ofattack_configdicts, one per chain step, using the same shape accepted by :meth:hack(each must include its ownattack_typeand any attack-specific settings). Only the first entry needs to specify how goals are sourced (goals,datasetorintents) unless thegoalsparameter below is provided; subsequent steps automatically receive only the goals still mitigated by the previous step (or all goals, seeescalate_only_mitigated). Defaults toNone, which resolves to the Jailbreak evaluation campaign's primary attacks, in order —h4rm3l→TAP→PAIR(seehackagent.risks.jailbreak.JAILBREAK_PROFILE). A goal source is still required either way, viagoalsor adataset/goals/intentskey on the first step.goals- Optional explicit list of goal strings to use for the whole chain. When provided, it takes precedence over anygoals/dataset/intentsset onattacks[0].run_config_override- Optional run configuration overrides applied to every step, forwarded to :meth:hack.fail_on_run_error- Forwarded to :meth:hackfor every step.escalate_only_mitigated- WhenTrue(default), a goal only moves on to the next attack if it was mitigated at the current step — goals that already succeeded are dropped, and each goal's final result is either its first success or its last (final) attempt. WhenFalse, every goal is sent to every attack regardless of outcome, and results from all steps are kept for every goal (nothing is dropped or overwritten).
Returns:
A flat list of result rows (same row shape as :meth:hack),
grouped by original goal, in first-seen order. Each row is
tagged with chain_step (0-based index into attacks) and
chain_attack_type identifying which attack produced it.
Raises:
HackAgentError- Ifattacksis empty, or a step is missingattack_type.