AI optimization loop

The AI optimization loop is not a structure to maximize the outcome of a single judgment. NoahAI is financial AI infrastructure designed so that, through the cycle of judgment → record → verify → feedback, judgment criteria themselves become more refined over time.

Today each account's records support review and safety controls for that account. A future cross-user layer remains conditional on explicit consent, anonymization, revocation, authorization, and operational validation.

1

Record

Judgment: structure decision support from market data

AI structures decision support from market data. Context and outcome of every judgment are recorded in a standardized format and stored for traceability.

2

Outcome

Outcome: record and explain results of judgment

Results of judgment are recorded and explained. The focus is not only on performance metrics or risk events but on clear explanation and record of outcomes.

3

Explain

Log: record judgment and outcome in explainable, standardized form

Judgment and outcome are fully recorded in explainable, standardized form. Under XAI policy, every decision process is transparent and categorized for traceability.

4

Policy

Replay: analyze logs and extract success/failure patterns

Recorded conditions, blocks, and outcomes are reviewed to derive improvement candidates. A result does not silently rewrite an existing policy; changes remain reviewable strategy or settings drafts.

5

Risk

Policy review: evaluate market- and asset-specific improvement candidates as a new version

Improvement candidates are evaluated separately by market regime and asset type. Strategy Studio changes require save, user approval, automated checks, and PAPER validation; LIVE policy is not changed without user confirmation.

6

Feedback

Feedback: detect risk signals and strengthen guardrails

Risk signals are detected early and conservative guardrails are strengthened when needed. In the current public release, feedback uses account-level records and review; execution remains optional and permission-bound. Anonymous cross-user patterns and shared-policy updates remain a roadmap pending consent, anonymization, revocation, and operational validation.

7

XAI

Explainable AI: explain every decision and maintain verifiable structure

Every decision's reasoning is kept explainable and audit logs are maintained. This step is essential for trust and transparency; local storage allows external verification.

Why the 'loop' matters in financial AI

Financial judgment depends on context: assets, liabilities, goals, living expenses, risk tolerance. Trust is built through repeated verification and feedback, not single outcomes. This structure can extend to voice phishing and fraud detection, protection of digitally vulnerable users, and other financial safety areas.

Operational view

The AI optimization loop is not for making more decisions; it is for reducing the chance of failure and gradually improving judgment criteria.

The 7-step loop operates as follows:

  • Reviewable cycle: The 7 steps connect records, explanations, outcomes, and versioned improvement candidates without silently rewriting existing policy.
  • Pattern-level learning: Learning is by success/failure patterns, not raw past performance, enabling regime-specific pattern learning.
  • Per-asset-type learning: Judgment context is separated by asset type; one asset's outcome does not directly affect other judgment domains.
  • Data-driven: All improvement is based on recorded data and outcomes; stability and reproducibility are validated in production.
  • Safety first: The Risk step prevents failure through conservative control and detects risk signals early.
  • Transparency: The XAI step keeps every decision's reasoning traceable; local storage allows external verification.
  • Current versus future: Account-level records support review now. Anonymous collective patterns are not yet an operating shared-policy feature.

Judgment-quality evaluation model

The formulas below are conceptual examples for explaining judgment quality. They are not a published production API, fixed weight contract, or claim of autonomous reinforcement learning, and they do not guarantee returns.

Actual evaluation keeps venue, market, strategy version, costs, data quality, guardrails, and PAPER/LIVE ledgers separate.

Profit trade reward

R_profit = α × profit_rate × confidence_score × (1 - risk_penalty)
  • α: Reward scaling (default 1.0)
  • profit_rate: Actual return (0.0–1.0)
  • confidence_score: AI confidence (0.0–1.0)
  • risk_penalty: Risk penalty (0.0–0.5)

Loss trade reward

R_loss = -β × |loss_rate| × (1 + consecutive_loss_penalty)
  • β: Loss scaling (default 1.2)
  • loss_rate: Actual loss (negative)
  • consecutive_loss_penalty: Consecutive loss penalty (0.0–0.3)

Risk management reward

R_risk_management = γ × (early_exit_bonus - late_exit_penalty)
  • γ: Risk management reward coefficient (default 0.5)
  • early_exit_bonus: Early exit bonus (0.0–0.2)
  • late_exit_penalty: Late exit penalty (0.0–0.3)

Supported paths are recorded in auditable form. Evaluation does not automatically change a strategy or LIVE setting; Strategy Studio changes require user approval, automatic checks, and PAPER validation.

How RL and collective learning connect

NoahAI's reinforcement learning does not maximize per-account returns. It rewards 'judgment quality' itself: appropriateness of judgment, risk response, explainability, accident avoidance.

Today each account's outcome supports that account's review and safety controls. A future cross-user layer would require explicit consent, anonymization, revocation, authorization, and operational validation before shared-policy updates. The design goal remains 'more users, lower probability of failure,' not an assertion that collective operation is complete.