System One (Decision) Models

Summary

System One Decision models API user guide. Models use AI but return structured decisions and output typed data or probabilities instead of generating paragraphs of text. Faster and cheaper than LLMs, these models can be integrated into workflows to route data and make decisions.

Body

Welcome to the Decision Router API. This service provides a unified FastAPI interface for high-volume policy checks, low-latency scoring, long-context log evaluations, and generative reasoning. These models can be used to quickly execute decisions where input is minimal and a typed data is desired. They can be integrated into workflows and applications in place of Large Language Models (LLMs) where speed and response structure are more important than generating text.

This API provides access to several open source decision model variants. A System One model is an artificial intelligence designed to make fast, structured decisions and output typed data or probabilities instead of generating paragraphs of text. The concept comes from psychologist Daniel Kahneman’s idea of "System 1" thinking, which is fast and automatic, compared to "System 2" thinking, which is slow and deliberate. Traditional chatbots act like System 2 by writing text token-by-token, while System One models are built for quick judgments that software can use immediately.

Advantages Limitations
Ultra-High Speed: Processes inputs in a single forward pass and skips token-by-token generation, running 100x to 200x faster than traditional LLMs. No Creative Output: Cannot generate prose, write essays, summarize articles in text, or compose emails.
Calibrated Probabilities: Outputs highly reliable percentages and scores (e.g., exactly how toxic or how relevant an input is). Fixed Output Formats: Completely restricted to strict typed questions, categories, or numbers.
Extremely Cost-Effective: Drastically lowers compute costs by eliminating heavy generative token overhead. No "CoT" Reasoning: Cannot use Chain-of-Thought reasoning to break down complex, multi-step math or logic problems before answering.Task-Specific: Built purely for evaluation, classification, routing, and guardrails, making it a poor choice for general-purpose chat.
Deterministic Parsing: Eliminates the flakiness of traditional LLMs, meaning software can reliably use its JSON data without parsing errors. Task-Specific: Built purely for evaluation, classification, routing, and guardrails, making it a poor choice for general-purpose chat.

Quick Start

Authorization X-API-Key
Base URL https://system-one.ai.hpc.fau.edu
List models /v1/models
Decide /v1/decide
Rotate your API Key

/v1/auth/rotate-key

 

Usage Examples

 

Rotate API Key

User Self-Service Key Rotation

If you suspect your API key has been exposed or need to rotate credentials according to your team's compliance schedule, you can generate a new key yourself.

Important: Rotating your key immediately revokes and invalidates your previous key.

 

Python Example

# python
import requests

API_URL = "https://system-one.ai.hpc.fau.edu"
current_key = "YOUR_CURRENT_KEY"

response = requests.post(
  f"{API_URL}/v1/auth/rotate-key",
  headers={"X-API-Key": current_key},
  )

if response.status_code == 200:
  new_key = response.json()["new_api_key"]
  print(f"Key rotated successfully. Your new key is: {new_key}")
else:
  print(f"Rotation failed: {response.json()}")

 

List Available Models

### Models endpoint

 

Check the models endpoint for exact model names and details.

 

# python

import requests

headers = {"X-API-Key": "YOUR_KEY"}

response = requests.get(
  "https://system-one.ai.hpc.fau.edu/v1/models",
  headers=headers,
  )

print(response.json())

 

Choosing the Right Engine Target (Model)

The /v1/decide endpoint accepts a model query parameter to route your request to the appropriate underlying CPU or GPU engine.

Engine Target Architecture Target Latency Max Context Best Used For
open-jev-deberta-v3-large CPU (ONNX) ~15 ms 512 tokens High-volume binary policy checks, guardrails, short log checks
laya CPU (ONNX) ~20 ms 512 tokens Confidence scoring, risk classification, escalation triggers
semif GPU (Gemma 4) ~80–150 ms 16,384 tokens Long-context documents, multi-option scoring, token-free decisions
simple-jev GPU (Gemma 4) ~300 ms–1.5 s 16,384 tokens Scenarios requiring free-form text explanations or detailed reasoning

Engine Model Examples

 

open-jev-deberta-v3-large

  • Best for: Real-time access control checks, fast policy verification, and low-latency microservice guardrails.
  • Limitations: Truncates inputs longer than 512 tokens. No narrative generation.

 

# python

import requests

headers = {"X-API-Key": "YOUR_KEY"}

payload = {
"state": "User accounts-service attempted 15 invalid login requests from IP 192.168.1.50 in 10 seconds.",
"questions": [
{
"instructions": "Should this IP address be rate-limited immediately?",
"type": "noul", # Yes/No evaluation
}
],
}

response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=open-jev-deberta-v3-large",
headers=headers,
json=payload,
)


print(response.json())
# Output: {"model_used": "open-jev-deberta-v3-large", "results": [{"p_yes": 0.9851, "p_no": 0.0149}]}

 

 

laya

  • Best for: Secondary validation and determining whether a human analyst needs to review a decision.
  • Limitations: Optimized for binary and triage logic; not intended for unstructured creative prompts.

 

# python

import requests

headers = {"X-API-Key": "YOUR_KEY"}

payload = {
"state": "High-value wire transfer requested for $45,000 from an unrecognized device in Paris.",
"questions": [
{
"instructions": "Does this transaction require manual Tier-2 fraud team escalation?",
"type": "noul",
}
],
}

response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=laya",
headers=headers,
json=payload,
)

 

semif

  • Best for: Analyzing long incident logs, long legal or compliance documents, or selecting the best option among multiple choices without the latency cost of generative LLM sampling.
  • Limitations: Returns option probabilities directly; it does not generate explanatory text paragraphs.

 

# python

import requests

headers = {"X-API-Key": "YOUR_KEY"}

payload = {
"state": """
[SYSTEM INCIDENT LOG]
08:12:01 - DB-01 CPU utilization exceeded 95%.
08:12:15 - Connection pool exhausted on auth-service-v2.
08:12:30 - Cache miss rate spikes on Redis node 3 following deployment commit #a8f3d1.
... [Multi-page log payload supported up to 16k tokens] ...

""",
"questions": [
{
"instructions": "Which subsystem is the primary root cause of this outage?",
"type": "choice",
"options": [
"Database CPU",
"Auth Service Memory",
"Redis Deployment",
"Network Partition",
],
}
],
}


response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=semif",
headers=headers,
json=payload,
)

print(response.json())
# Output: {"model_used": "semif", "results": [{"top_option": "Redis Deployment", "probabilities": {"Redis Deployment": 0.912, ...}}]}

 

simple-jev

  • Best for: Audit trails, natural-language explanations for decisions, and complex reasoning tasks.
  • Limitations: Higher latency due to token-by-token generation.

 

# python

import requests

headers = {"X-API-Key": "YOUR_KEY"}

payload = {
"state": "Customer requested a refund 35 days after purchase. Standard policy is 30 days, but the customer's product arrived damaged.",
"questions": [
{
"instructions": "Provide a policy exception decision and explain the rationale.",
"type": "choice",
"options": ["Approve Refund Exception", "Deny Refund"],
}
],
}

response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=simple-jev",
headers=headers,
json=payload,
)

 

 

 

API Error Handling Reference

 

Status Code Description Resolution

401 Unauthorized

Missing or revokedX-API-Key

Verify the header string or rotate the key through /v1/auth/rotate-key.

400 Bad Request Missing required parameters, such as `options` for `type: "choice"` Check the request payload against the Swagger API schemas (https://system-one.ai.hpc.fau.edu/docs).
500 Server Error Model engine inference timeout or backend error Check the API health status (https://system-one.ai.hpc.fau.edu/health). Contact hpc@fau.edu if down.

 

Details

Details

Article ID: 164964
Created
Thu 10/1/26 4:37 PM
Modified
Thu 10/1/26 5:10 PM