Body
Welcome to the Decision Router API. This service provides a unified FastAPI interface for high-volume policy checks, low-latency scoring, long-context log evaluations, and generative reasoning. These models can be used to quickly execute decisions where input is minimal and a typed data is desired. They can be integrated into workflows and applications in place of Large Language Models (LLMs) where speed and response structure are more important than generating text.
This API provides access to several open source decision model variants. A System One model is an artificial intelligence designed to make fast, structured decisions and output typed data or probabilities instead of generating paragraphs of text. The concept comes from psychologist Daniel Kahneman’s idea of "System 1" thinking, which is fast and automatic, compared to "System 2" thinking, which is slow and deliberate. Traditional chatbots act like System 2 by writing text token-by-token, while System One models are built for quick judgments that software can use immediately.
| Advantages |
Limitations |
| Ultra-High Speed: Processes inputs in a single forward pass and skips token-by-token generation, running 100x to 200x faster than traditional LLMs. |
No Creative Output: Cannot generate prose, write essays, summarize articles in text, or compose emails. |
| Calibrated Probabilities: Outputs highly reliable percentages and scores (e.g., exactly how toxic or how relevant an input is). |
Fixed Output Formats: Completely restricted to strict typed questions, categories, or numbers. |
| Extremely Cost-Effective: Drastically lowers compute costs by eliminating heavy generative token overhead. |
No "CoT" Reasoning: Cannot use Chain-of-Thought reasoning to break down complex, multi-step math or logic problems before answering.Task-Specific: Built purely for evaluation, classification, routing, and guardrails, making it a poor choice for general-purpose chat. |
| Deterministic Parsing: Eliminates the flakiness of traditional LLMs, meaning software can reliably use its JSON data without parsing errors. |
Task-Specific: Built purely for evaluation, classification, routing, and guardrails, making it a poor choice for general-purpose chat. |
Quick Start
Usage Examples
Rotate API Key
User Self-Service Key Rotation
If you suspect your API key has been exposed or need to rotate credentials according to your team's compliance schedule, you can generate a new key yourself.
Important: Rotating your key immediately revokes and invalidates your previous key.
Python Example
# python
import requests
API_URL = "https://system-one.ai.hpc.fau.edu"
current_key = "YOUR_CURRENT_KEY"
response = requests.post(
f"{API_URL}/v1/auth/rotate-key",
headers={"X-API-Key": current_key},
)
if response.status_code == 200:
new_key = response.json()["new_api_key"]
print(f"Key rotated successfully. Your new key is: {new_key}")
else:
print(f"Rotation failed: {response.json()}")
List Available Models
### Models endpoint
Check the models endpoint for exact model names and details.
# python
import requests
headers = {"X-API-Key": "YOUR_KEY"}
response = requests.get(
"https://system-one.ai.hpc.fau.edu/v1/models",
headers=headers,
)
print(response.json())
Choosing the Right Engine Target (Model)
The /v1/decide endpoint accepts a model query parameter to route your request to the appropriate underlying CPU or GPU engine.
| Engine Target |
Architecture |
Target Latency |
Max Context |
Best Used For |
| open-jev-deberta-v3-large |
CPU (ONNX) |
~15 ms |
512 tokens |
High-volume binary policy checks, guardrails, short log checks |
| laya |
CPU (ONNX) |
~20 ms |
512 tokens |
Confidence scoring, risk classification, escalation triggers |
| semif |
GPU (Gemma 4) |
~80–150 ms |
16,384 tokens |
Long-context documents, multi-option scoring, token-free decisions |
| simple-jev |
GPU (Gemma 4) |
~300 ms–1.5 s |
16,384 tokens |
Scenarios requiring free-form text explanations or detailed reasoning |
Engine Model Examples
open-jev-deberta-v3-large
- Best for: Real-time access control checks, fast policy verification, and low-latency microservice guardrails.
- Limitations: Truncates inputs longer than 512 tokens. No narrative generation.
# python
import requests
headers = {"X-API-Key": "YOUR_KEY"}
payload = {
"state": "User accounts-service attempted 15 invalid login requests from IP 192.168.1.50 in 10 seconds.",
"questions": [
{
"instructions": "Should this IP address be rate-limited immediately?",
"type": "noul", # Yes/No evaluation
}
],
}
response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=open-jev-deberta-v3-large",
headers=headers,
json=payload,
)
print(response.json())
# Output: {"model_used": "open-jev-deberta-v3-large", "results": [{"p_yes": 0.9851, "p_no": 0.0149}]}
laya
- Best for: Secondary validation and determining whether a human analyst needs to review a decision.
- Limitations: Optimized for binary and triage logic; not intended for unstructured creative prompts.
# python
import requests
headers = {"X-API-Key": "YOUR_KEY"}
payload = {
"state": "High-value wire transfer requested for $45,000 from an unrecognized device in Paris.",
"questions": [
{
"instructions": "Does this transaction require manual Tier-2 fraud team escalation?",
"type": "noul",
}
],
}
response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=laya",
headers=headers,
json=payload,
)
semif
- Best for: Analyzing long incident logs, long legal or compliance documents, or selecting the best option among multiple choices without the latency cost of generative LLM sampling.
- Limitations: Returns option probabilities directly; it does not generate explanatory text paragraphs.
# python
import requests
headers = {"X-API-Key": "YOUR_KEY"}
payload = {
"state": """
[SYSTEM INCIDENT LOG]
08:12:01 - DB-01 CPU utilization exceeded 95%.
08:12:15 - Connection pool exhausted on auth-service-v2.
08:12:30 - Cache miss rate spikes on Redis node 3 following deployment commit #a8f3d1.
... [Multi-page log payload supported up to 16k tokens] ...
""",
"questions": [
{
"instructions": "Which subsystem is the primary root cause of this outage?",
"type": "choice",
"options": [
"Database CPU",
"Auth Service Memory",
"Redis Deployment",
"Network Partition",
],
}
],
}
response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=semif",
headers=headers,
json=payload,
)
print(response.json())
# Output: {"model_used": "semif", "results": [{"top_option": "Redis Deployment", "probabilities": {"Redis Deployment": 0.912, ...}}]}
simple-jev
- Best for: Audit trails, natural-language explanations for decisions, and complex reasoning tasks.
- Limitations: Higher latency due to token-by-token generation.
# python
import requests
headers = {"X-API-Key": "YOUR_KEY"}
payload = {
"state": "Customer requested a refund 35 days after purchase. Standard policy is 30 days, but the customer's product arrived damaged.",
"questions": [
{
"instructions": "Provide a policy exception decision and explain the rationale.",
"type": "choice",
"options": ["Approve Refund Exception", "Deny Refund"],
}
],
}
response = requests.post(
"https://system-one.ai.hpc.fau.edu/v1/decide?model=simple-jev",
headers=headers,
json=payload,
)
API Error Handling Reference
| Status Code |
Description |
Resolution |
|
401 Unauthorized
|
Missing or revokedX-API-Key
|
Verify the header string or rotate the key through /v1/auth/rotate-key.
|
| 400 Bad Request |
Missing required parameters, such as `options` for `type: "choice"` |
Check the request payload against the Swagger API schemas (https://system-one.ai.hpc.fau.edu/docs). |
| 500 Server Error |
Model engine inference timeout or backend error |
Check the API health status (https://system-one.ai.hpc.fau.edu/health). Contact hpc@fau.edu if down. |