AI-Native Research Infrastructure

How Agents Work at Arc Research

A technical deep-dive into the architecture that powers our commodity intelligence platform. From vector embeddings and knowledge graphs to tool-augmented LLM agents, here is exactly how we turn raw data into actionable research.

The 30-Second Version

TL;DR for the Time-Pressed

Arc Research is an AI-native research platform purpose-built for commodity traders. Every document you upload, every feed you subscribe to, every note you write gets chunked, embedded, and indexed into a personal knowledge graph. When you chat with our agents, they don't just answer from their training data — they search your documents, pull live futures quotes, reference your research stories, and write journal entries — all in real-time. Archie Desk is the paper-trading layer: playbooks, review quorum, and inspectable fills when a detector fires.

Think of it as a second brain with a Bloomberg Terminal's memory, an analyst's judgment, and a paper blotter you can audit.

System Architecture

The Full Pipeline: From Data to Insight

When you interact with Arc Research, you're touching a multi-layered system that ingests, processes, indexes, and retrieves information across several dimensions simultaneously. Here's the high-level architecture:

flowchart TB
    subgraph Ingestion["Data Ingestion Layer"]
        A["PDFs & Documents"] --> C["Text Extraction"]
        B["Feeds: CFTC, EIA,
Podcasts, Newsletters"] --> D["Feed Processors"] E["User Notes &
Journal Entries"] --> F["Direct Input"] end subgraph Processing["Processing Pipeline"] C --> G["Recursive Character
Chunking"] D --> H["Feed-Specific
Chunking"] F --> I["Context Encoding"] G --> J["Embedding Service
OpenAI text-embedding-3-small"] H --> J I --> J end subgraph Storage["Knowledge Store"] J --> K[("pgvector
1536-dim Vectors")] J --> L[("Apache AGE
Knowledge Graph")] K --> M["Document Chunks"] K --> N["Feed Chunks"] L --> O["Entities &
Relationships"] end subgraph Agents["Agent Layer"] P["User Query"] --> Q["Chat Agent"] Q --> R{"RAG Retrieval"} R --> K Q --> S["Tool Registry"] S --> T["Futures Quotes"] S --> U["Document Search"] S --> V["Story Management"] S --> W["Document Analysis"] AA["Market Event"] --> AB["Archie Desk"] AB --> AC["Playbook + Quorum"] AC --> AD["Paper Book"] end R --> X["Context-Enriched
Response"] S --> X X --> Y["Streamed to User
via Turbo Streams"] style Ingestion fill:#1e1b4b,stroke:#818cf8,color:#ffffff style Processing fill:#312e81,stroke:#818cf8,color:#ffffff style Storage fill:#1e1b4b,stroke:#818cf8,color:#ffffff style Agents fill:#312e81,stroke:#818cf8,color:#ffffff

End-to-end architecture: from raw data ingestion to streamed agent responses and a paper desk.

Archie Desk

From Signal to Inspectable Paper Decision

Chat is for questions. The desk is for decisions. When a golden cross, death cross, COT update, EIA print, price alert, or watcher flag fires, Archie loads a playbook and the current book snapshot, then proposes enter, skip, hold, tighten, or exit.

Review playbooks vote. Hard stops and time stops are enforced in code. Fills land on a paper book you can open later — rationale, votes, and tool calls included. Archie does not send unsupervised live orders.

sequenceDiagram
    participant Detector as Detector / Feed
    participant Wake as WakeService
    participant Archie as Desk Supervisor
    participant Review as Review Quorum
    participant Book as Paper Book

    Detector->>Wake: golden_cross / cot_flip / eia_surprise
    Wake->>Wake: Position awareness
    Wake->>Archie: Entry or monitor playbook
    Archie->>Archie: propose_decision
    Archie->>Review: Regime + positioning + news
    Review-->>Book: Quorum pass → paper fill
    Review-->>Archie: Veto or fail → rejected
          

Event in, playbook out, paper fill only after review. Policy stops can exit without a model vote.

Retrieval-Augmented Generation

How RAG Makes Agents Actually Useful

If you've used ChatGPT, you know the problem: LLMs are trained on static data and have no idea what's in your research pipeline. RAG (Retrieval-Augmented Generation) fixes this by injecting relevant context into the prompt at query time.

ELI5: How RAG Works

Imagine you're a trader and you ask your analyst: "What's the latest CFTC positioning on natural gas?" A dumb analyst would guess from memory. A smart analyst would first pull the latest report from the filing cabinet, read the relevant sections, and then give you an informed answer. That's RAG.

  1. Embed the question — Your query is converted into a 1,536-dimensional vector using OpenAI's text-embedding-3-small model. This vector captures the semantic meaning, not just keywords.
  2. Search your knowledge base — We run a cosine similarity search against every document chunk and feed chunk you've indexed in pgvector. The top 5 most semantically relevant chunks are retrieved — across all your source documents and feed data.
  3. Inject into the prompt — Those chunks become system-level context for the LLM. The agent now has specific, factual material from your data to reference when generating a response.
  4. Generate with grounding — The LLM composes its response using both its general training and your specific research context. It cites your documents and attributes information to specific sources.
sequenceDiagram
    actor Trader
    participant Chat as Chat Interface
    participant Embed as Embedding Service
    participant PGV as pgvector DB
    participant LLM as LLM Agent
    participant Tools as Tool Registry

    Trader->>Chat: "What's the latest spec positioning in crude?"
    Chat->>Embed: Embed query → 1536-dim vector
    Embed->>PGV: Cosine similarity search
    PGV-->>Chat: Top 5 relevant chunks
    Note over PGV,Chat: CFTC feed chunks, research docs,
journal entries Chat->>LLM: System prompt + RAG context + query LLM->>Tools: call FuturesQuoteTool(CL) Tools-->>LLM: CL1! = $72.45 (+0.8%) LLM->>Tools: call SearchDocumentsTool("crude positioning") Tools-->>LLM: 3 additional document matches LLM-->>Chat: Streamed response with citations Chat-->>Trader: Real-time Turbo Stream update

Sequence diagram: A single chat query triggers embedding, retrieval, tool calls, and streaming.

Document Intelligence

From PDF to Searchable Knowledge

When you upload a document — a broker research PDF, a quarterly earnings transcript, your own position notes — it doesn't just sit in a folder. It gets immediately processed through a multi-stage pipeline:

flowchart LR
    A["Upload
PDF / DOCX / TXT"] --> B["Text Extraction
pdf-reader / docx gem"] B --> C["Metadata Parsing
Gemini 2.0 Flash"] B --> D["Recursive Character
Chunking"] D --> E["Chunk 1
≤1000 chars"] D --> F["Chunk 2
≤1000 chars"] D --> G["Chunk N
≤1000 chars"] E --> H["Concurrent Embedding
10 parallel fibers"] F --> H G --> H H --> I[("pgvector
1536-dim per chunk")] C --> J["Structured Metadata
title, author, keywords"] style A fill:#4338ca,stroke:#a5b4fc,color:#ffffff style I fill:#312e81,stroke:#818cf8,color:#ffffff style J fill:#312e81,stroke:#818cf8,color:#ffffff

Recursive Character Chunking

Documents are split using a hierarchy of separators: paragraph breaks first, then line breaks, then spaces. Each chunk is capped at 1,000 characters with 100-character overlap between adjacent chunks. This preserves sentence-level coherence while keeping chunks small enough for precise retrieval.

Concurrent Embedding

We use Ruby's Async fiber scheduler to embed 10 chunks in parallel. A 50-page research PDF generates ~200 chunks and is fully indexed in seconds, not minutes. Each chunk becomes a 1,536-dimensional vector via OpenAI's embedding model.

AI Metadata Extraction

On upload, Gemini 2.0 Flash parses the first section of each document to extract structured metadata: author, title, description, publication, keywords, and dates. This metadata enriches search results and feeds into the knowledge graph.

User-Scoped Privacy

Every embedding is scoped to your user account. When the RAG pipeline searches for relevant context, it only queries your document chunks. Your proprietary research and notes are never exposed to other users or leaked into shared model context.

Tool-Augmented Agents

Agents That Do Things, Not Just Talk

The real power isn't just in answering questions — it's in taking action. Our agents have access to a registry of tools they can invoke mid-conversation. When you ask a question that requires live data, document lookups, or write operations, the agent autonomously decides which tools to call, in what order, and how to synthesize the results.

This is function calling at the LLM layer. The model doesn't execute code — it emits structured JSON describing which tool to call and with what parameters. Our server-side tool registry handles execution and returns results back into the conversation context.

FuturesQuoteTool

Pulls real-time futures quotes from the Polygon API. Ask "What's gold trading at?" and the agent calls execute(symbol: "GC") to fetch the latest price, change, and volume. No stale training data — live market data streamed into the conversation.

SearchDocumentsTool

Performs semantic search across your entire document library. Uses the same embedding + cosine similarity pipeline as RAG, but invoked on-demand by the agent. Useful when the initial RAG context isn't enough and the agent needs to dig deeper.

DocumentAnalysisTool

Generates summaries, outlines, and concept maps from any source document. Results are cached so the first analysis is LLM-powered, and subsequent requests are instant. Ask "Summarize that Goldman Sachs copper report" and the agent handles the rest.

AddJournalEntryTool & ListStoriesTool

Agents can read and write to your research stories. Mid-conversation, you can say "Add this analysis to my Nat Gas Winter Thesis story" and the agent creates a timestamped journal entry. ListStoriesTool lets it browse your full story inventory with associated tickers and recent entries.

ChatContextInjectionTool

Resolves @mentions in your messages. Reference a ticker, story, position, or company by name and the tool injects the full context — current prices, thesis details, P&L, fundamentals — directly into the conversation. It's how agents "see" the entities you're talking about.

Multi-Model Orchestration

Right Model for the Right Job

We don't believe in a one-model-fits-all approach. Different tasks have fundamentally different cost-latency-quality tradeoffs, and we route accordingly. Here's how we allocate model capacity:

Task Model Why
Interactive Chat gemini-2.5-flash Fast streaming, good tool use, low cost per token
CFTC & EIA Report Generation gemini-2.5-pro Complex tabular reasoning, long-context analysis
Story Ticker Enrichment gpt-4o-mini High throughput, good at structured reasoning
Metadata Extraction gemini-2.5-flash Schema-constrained output, fast, cost-efficient
Embeddings text-embedding-3-small 1536 dims, best cost/quality for semantic search
Graph Visualization gemini-2.5-flash Creative structured output, Mermaid syntax generation

We use RubyLLM as our abstraction layer across providers. This gives us a unified interface for chat, embeddings, and tool use across OpenAI and Google Gemini, with the ability to swap or add providers without changing application code. Model metadata — context windows, capabilities, pricing — is tracked in our database so we can route dynamically.

Automated Feed Intelligence

Government Data to Trading Intelligence in Minutes

The CFTC Commitments of Traders report and the EIA Natural Gas Storage report are two of the most watched data releases in commodity markets. We've built automated pipelines that fetch, parse, embed, and analyze these reports the moment they drop.

flowchart TB
    subgraph Schedule["Scheduled Triggers"]
        A["CFTC: Fridays @ 2:30 PM ET"]
        B["EIA NG: Thursdays @ 10:30 AM ET"]
    end

    subgraph Fetch["Fetch & Transform"]
        A --> C["FetchCFTCReportsJob"]
        B --> D["FetchEIANaturalGas
StorageReportsJob"] C --> E["Create FeedRun
state: processing"] D --> E E --> F["fetch_data → raw"] F --> G["transform_data → normalized"] G --> H["validate_schema"] H --> I["Store result_data
as JSONB"] end subgraph Embed["Embed & Index"] I --> J["GenerateFeedEmbeddingsJob"] J --> K["Feed-Specific
Chunking Strategy"] K --> L["Concurrent Embedding
10 fibers"] L --> M[("FeedChunk records
in pgvector")] end subgraph Analyze["AI Report Generation"] I --> N["GenerateCFTCReport
or GenerateEIAReport"] N --> O["Gemini 2.5 Pro
with specialized prompt"] O --> P["Markdown Report
with analysis"] P --> Q["Published to
Reports feed"] end style Schedule fill:#1e1b4b,stroke:#818cf8,color:#ffffff style Fetch fill:#312e81,stroke:#818cf8,color:#ffffff style Embed fill:#1e1b4b,stroke:#818cf8,color:#ffffff style Analyze fill:#312e81,stroke:#818cf8,color:#ffffff

Automated pipeline: scheduled fetch, schema validation, concurrent embedding, and AI report generation.

What This Means for You

By the time you sit down Monday morning, last Friday's CFTC data has already been:

  • Fetched and validated against the expected schema
  • Chunked and embedded into your searchable knowledge base
  • Analyzed by Gemini 2.5 Pro into a full markdown report with positioning changes, open interest trends, and notable shifts
  • Made available as RAG context for your next chat conversation

No manual parsing. No spreadsheet wrangling. The intelligence is there when you need it.

Knowledge Graph

Beyond Vector Search: Structured Relationships

Vector embeddings are powerful for semantic similarity, but they don't capture relationships. Knowing that a document chunk about "copper mine expansions" is semantically close to "Freeport-McMoRan" is useful. Knowing that Freeport-McMoRan operates the Grasberg mine, which produces copper and gold, and that the CEO discussed capex guidance on the Q3 call — that's a knowledge graph.

We use Apache AGE (A Graph Extension) running on PostgreSQL to store entity-relationship triples alongside our vector store. This means a single query can combine:

  • Semantic search (pgvector) — "find chunks about copper supply disruptions"
  • Graph traversal (Apache AGE) — "which companies are connected to those disruptions?"
  • Structured data (PostgreSQL) — "what are those companies' latest futures positions?"
graph LR
    A(("Freeport-
McMoRan")) -->|OPERATES| B(("Grasberg
Mine")) B -->|PRODUCES| C(("Copper")) B -->|PRODUCES| D(("Gold")) A -->|CEO| E(("Kathleen
Quirk")) E -->|DISCUSSED| F(("Q3 2025
Earnings Call")) F -->|MENTIONS| G(("Capex
Guidance")) F -->|MENTIONS| C C -->|TRACKED_BY| H(("CFTC COT
Report")) H -->|SHOWS| I(("Spec Net Long
+42k contracts")) style A fill:#4338ca,stroke:#a5b4fc,color:#ffffff style B fill:#312e81,stroke:#818cf8,color:#ffffff style C fill:#b45309,stroke:#fbbf24,color:#ffffff style D fill:#a16207,stroke:#fbbf24,color:#ffffff style E fill:#4338ca,stroke:#a5b4fc,color:#ffffff style F fill:#312e81,stroke:#818cf8,color:#ffffff style G fill:#312e81,stroke:#818cf8,color:#ffffff style H fill:#1e3a5f,stroke:#60a5fa,color:#ffffff style I fill:#1e3a5f,stroke:#60a5fa,color:#ffffff

Example knowledge graph: entities, relationships, and data sources connected across your research.

Graph Node Types

When our agents process documents, they extract structured entities into the graph. Typical node labels include:

PERSON

Analysts, executives, officials

ORGANIZATION

Companies, exchanges, regulators

CONCEPT

Contango, backwardation, carry

TOPIC

Supply disruption, weather, geopolitics

Real-Time Delivery

Streaming Architecture: No Waiting for Full Responses

When you send a message, you don't wait 30 seconds for a complete response. Tokens start appearing in your chat window within milliseconds of the first LLM output. Here's how:

sequenceDiagram
    actor User
    participant Controller as Rails Controller
    participant Sidekiq as Sidekiq Job
    participant LLM as LLM Provider
    participant Turbo as Turbo Streams

    User->>Controller: POST /chat_messages
    Controller->>Controller: Create ChatMessage record
    Controller->>Sidekiq: ChatStreamJob.perform_later(chat_id)
    Controller-->>User: Turbo Stream: append message shell

    Sidekiq->>Sidekiq: Register tools from ToolRegistry
    Sidekiq->>LLM: chat.complete (streaming)

    loop Token by token
        LLM-->>Sidekiq: chunk
        Sidekiq->>Turbo: broadcast_append_chunk
        Turbo-->>User: Real-time DOM update
    end

    Note over Sidekiq,LLM: Tool calls happen mid-stream:
Agent pauses, calls tool, resumes generation Sidekiq->>Sidekiq: Save final assistant message

Sidekiq

Chat streaming runs in dedicated background workers. Your browser isn't blocking on LLM inference.

Turbo Streams

Hotwire's Turbo Streams push DOM updates over WebSocket. Each token chunk triggers a targeted HTML append.

Zero JS Bloat

No React. No client-side state management. Server-rendered HTML streamed in real-time via Rails conventions.

Context Engine

Every Entity Is a First-Class Citizen

In Arc Research, tickers, stories, positions, companies, and users are all "mentionable" entities. When you type @GC (gold futures) or @My Copper Thesis in a chat message, the ChatContextInjectionService resolves those references and injects the full entity context into the conversation.

This means the agent isn't just reading text — it knows what you're talking about. If you mention a story, the agent sees the story's title, associated tickers, and the latest journal entry. If you mention a ticker, it gets the current quote, your position (if any), and relevant research.

flowchart LR
    A["User message:
'How does @GC positioning
look vs my @Gold Thesis?'"] --> B["Context Injection
Service"] B --> C["Resolve @GC
→ Ticker context"] B --> D["Resolve @Gold Thesis
→ Story context"] C --> E["GC1! $2,340.50
Vol: 234k | OI: 512k"] D --> F["Story: Gold Thesis
Tickers: GC, GLD, NEM
Last entry: Feb 12"] E --> G["Enriched Prompt
to LLM Agent"] F --> G G --> H["Agent response with
full entity awareness"] style A fill:#312e81,stroke:#818cf8,color:#ffffff style B fill:#4338ca,stroke:#a5b4fc,color:#ffffff style C fill:#4338ca,stroke:#a5b4fc,color:#ffffff style D fill:#4338ca,stroke:#a5b4fc,color:#ffffff style E fill:#1e3a5f,stroke:#60a5fa,color:#ffffff style F fill:#1e3a5f,stroke:#60a5fa,color:#ffffff style G fill:#4338ca,stroke:#a5b4fc,color:#ffffff style H fill:#312e81,stroke:#818cf8,color:#ffffff

Under the Hood

The Infrastructure Stack

We built this on battle-tested technology. No experimental frameworks. No hype-driven decisions. Every component was chosen because it's the right tool for the job in a latency-sensitive, data-intensive research application.

Application

  • Ruby on Rails 7 — Server-rendered, full-stack
  • Hotwire (Turbo + Stimulus) — Real-time UI without SPA complexity
  • Tailwind CSS — Utility-first styling
  • Sorbet — Gradual static typing for Ruby

Data & AI

  • PostgreSQL + pgvector — Relational + vector in one DB
  • Apache AGE — Graph queries on PostgreSQL
  • RubyLLM — Multi-provider LLM abstraction
  • Sidekiq — Background job processing for AI workloads

LLM Providers

  • Google Gemini — Chat, analysis, report generation
  • OpenAI — Embeddings, enrichment tasks

Data Sources

  • Polygon.io — Real-time futures quotes
  • CFTC — Commitments of Traders data
  • EIA — Energy storage and production data
  • Your uploads — PDFs, DOCX, notes, research

The Bottom Line

Why This Architecture Matters for Trading

You could use a general-purpose AI chatbot. You could copy-paste CFTC data into a spreadsheet. You could manually search through PDFs and broker notes. Traders have been doing these things for years. Here's why this is different:

1

Compounding Context

Every document you add, every feed that runs, every journal entry you write makes the system smarter for your specific research workflow. It's not starting from zero each conversation — it's building on everything you've fed it. Think of it as compound interest for your research graph.

2

Cross-Source Synthesis

When you ask about copper positioning, the agent can simultaneously reference your uploaded Goldman Sachs research PDF, the latest CFTC feed data, your story about Freeport-McMoRan, and the live CL quote — all in one response. No human analyst can context-switch across that many sources that fast.

3

Always-On Data Processing

CFTC data hits at 2:30 PM Friday. EIA storage at 10:30 AM Thursday. Our background jobs are already running when the data drops — fetching, validating, embedding, and generating analysis. By the time you check in, the work is done.

4

Institutional-Grade Privacy

Your research is your edge. Every vector embedding is user-scoped. Your documents, your stories, your positions never leak into another user's context. The knowledge graph is yours alone.

5

MCP in your editor

With Arc Research Starter, the same knowledge graph and market tools are available over MCP in Cursor and Claude. Generate a token, connect once, and keep researching where you already work. See MCP setup.

6

A paper desk with a paper trail

When a detector fires, Archie can propose a paper trade, send it through review, and leave a run you can audit. Hard stops stay in code. Live tickets stay gated. Open Archie Desk.

Ready to Put AI to Work on Your Research?

Start building your personal knowledge graph. Upload a document, chat with an agent, connect MCP from Cursor, or open Archie Desk and read the next paper decision.