AI-Native Research Infrastructure
A technical deep-dive into the architecture that powers our commodity intelligence platform. From vector embeddings and knowledge graphs to tool-augmented LLM agents, here is exactly how we turn raw data into actionable research.
TL;DR for the Time-Pressed
Arc Research is an AI-native research platform purpose-built for commodity traders. Every document you upload, every feed you subscribe to, every note you write gets chunked, embedded, and indexed into a personal knowledge graph. When you chat with our agents, they don't just answer from their training data — they search your documents, pull live futures quotes, reference your research stories, and write journal entries — all in real-time. Archie Desk is the paper-trading layer: playbooks, review quorum, and inspectable fills when a detector fires.
Think of it as a second brain with a Bloomberg Terminal's memory, an analyst's judgment, and a paper blotter you can audit.
The Full Pipeline: From Data to Insight
When you interact with Arc Research, you're touching a multi-layered system that ingests, processes, indexes, and retrieves information across several dimensions simultaneously. Here's the high-level architecture:
flowchart TB
subgraph Ingestion["Data Ingestion Layer"]
A["PDFs & Documents"] --> C["Text Extraction"]
B["Feeds: CFTC, EIA,
Podcasts, Newsletters"] --> D["Feed Processors"]
E["User Notes &
Journal Entries"] --> F["Direct Input"]
end
subgraph Processing["Processing Pipeline"]
C --> G["Recursive Character
Chunking"]
D --> H["Feed-Specific
Chunking"]
F --> I["Context Encoding"]
G --> J["Embedding Service
OpenAI text-embedding-3-small"]
H --> J
I --> J
end
subgraph Storage["Knowledge Store"]
J --> K[("pgvector
1536-dim Vectors")]
J --> L[("Apache AGE
Knowledge Graph")]
K --> M["Document Chunks"]
K --> N["Feed Chunks"]
L --> O["Entities &
Relationships"]
end
subgraph Agents["Agent Layer"]
P["User Query"] --> Q["Chat Agent"]
Q --> R{"RAG Retrieval"}
R --> K
Q --> S["Tool Registry"]
S --> T["Futures Quotes"]
S --> U["Document Search"]
S --> V["Story Management"]
S --> W["Document Analysis"]
AA["Market Event"] --> AB["Archie Desk"]
AB --> AC["Playbook + Quorum"]
AC --> AD["Paper Book"]
end
R --> X["Context-Enriched
Response"]
S --> X
X --> Y["Streamed to User
via Turbo Streams"]
style Ingestion fill:#1e1b4b,stroke:#818cf8,color:#ffffff
style Processing fill:#312e81,stroke:#818cf8,color:#ffffff
style Storage fill:#1e1b4b,stroke:#818cf8,color:#ffffff
style Agents fill:#312e81,stroke:#818cf8,color:#ffffff
End-to-end architecture: from raw data ingestion to streamed agent responses and a paper desk.
From Signal to Inspectable Paper Decision
Chat is for questions. The desk is for decisions. When a golden cross, death cross, COT update, EIA print, price alert, or watcher flag fires, Archie loads a playbook and the current book snapshot, then proposes enter, skip, hold, tighten, or exit.
Review playbooks vote. Hard stops and time stops are enforced in code. Fills land on a paper book you can open later — rationale, votes, and tool calls included. Archie does not send unsupervised live orders.
sequenceDiagram
participant Detector as Detector / Feed
participant Wake as WakeService
participant Archie as Desk Supervisor
participant Review as Review Quorum
participant Book as Paper Book
Detector->>Wake: golden_cross / cot_flip / eia_surprise
Wake->>Wake: Position awareness
Wake->>Archie: Entry or monitor playbook
Archie->>Archie: propose_decision
Archie->>Review: Regime + positioning + news
Review-->>Book: Quorum pass → paper fill
Review-->>Archie: Veto or fail → rejected
Event in, playbook out, paper fill only after review. Policy stops can exit without a model vote.
How RAG Makes Agents Actually Useful
If you've used ChatGPT, you know the problem: LLMs are trained on static data and have no idea what's in your research pipeline. RAG (Retrieval-Augmented Generation) fixes this by injecting relevant context into the prompt at query time.
Imagine you're a trader and you ask your analyst: "What's the latest CFTC positioning on natural gas?" A dumb analyst would guess from memory. A smart analyst would first pull the latest report from the filing cabinet, read the relevant sections, and then give you an informed answer. That's RAG.
text-embedding-3-small model.
This vector captures the semantic meaning, not just keywords.
pgvector. The top 5 most semantically relevant
chunks are retrieved — across all your source documents and feed data.
sequenceDiagram
actor Trader
participant Chat as Chat Interface
participant Embed as Embedding Service
participant PGV as pgvector DB
participant LLM as LLM Agent
participant Tools as Tool Registry
Trader->>Chat: "What's the latest spec positioning in crude?"
Chat->>Embed: Embed query → 1536-dim vector
Embed->>PGV: Cosine similarity search
PGV-->>Chat: Top 5 relevant chunks
Note over PGV,Chat: CFTC feed chunks, research docs,
journal entries
Chat->>LLM: System prompt + RAG context + query
LLM->>Tools: call FuturesQuoteTool(CL)
Tools-->>LLM: CL1! = $72.45 (+0.8%)
LLM->>Tools: call SearchDocumentsTool("crude positioning")
Tools-->>LLM: 3 additional document matches
LLM-->>Chat: Streamed response with citations
Chat-->>Trader: Real-time Turbo Stream update
Sequence diagram: A single chat query triggers embedding, retrieval, tool calls, and streaming.
From PDF to Searchable Knowledge
When you upload a document — a broker research PDF, a quarterly earnings transcript, your own position notes — it doesn't just sit in a folder. It gets immediately processed through a multi-stage pipeline:
flowchart LR
A["Upload
PDF / DOCX / TXT"] --> B["Text Extraction
pdf-reader / docx gem"]
B --> C["Metadata Parsing
Gemini 2.0 Flash"]
B --> D["Recursive Character
Chunking"]
D --> E["Chunk 1
≤1000 chars"]
D --> F["Chunk 2
≤1000 chars"]
D --> G["Chunk N
≤1000 chars"]
E --> H["Concurrent Embedding
10 parallel fibers"]
F --> H
G --> H
H --> I[("pgvector
1536-dim per chunk")]
C --> J["Structured Metadata
title, author, keywords"]
style A fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style I fill:#312e81,stroke:#818cf8,color:#ffffff
style J fill:#312e81,stroke:#818cf8,color:#ffffff
Documents are split using a hierarchy of separators: paragraph breaks first, then line breaks, then spaces. Each chunk is capped at 1,000 characters with 100-character overlap between adjacent chunks. This preserves sentence-level coherence while keeping chunks small enough for precise retrieval.
We use Ruby's Async fiber scheduler to embed
10 chunks in parallel. A 50-page research PDF generates ~200 chunks and is fully indexed in seconds,
not minutes. Each chunk becomes a 1,536-dimensional vector via OpenAI's embedding model.
On upload, Gemini 2.0 Flash parses the first section of each document to extract structured metadata: author, title, description, publication, keywords, and dates. This metadata enriches search results and feeds into the knowledge graph.
Every embedding is scoped to your user account. When the RAG pipeline searches for relevant context, it only queries your document chunks. Your proprietary research and notes are never exposed to other users or leaked into shared model context.
Agents That Do Things, Not Just Talk
The real power isn't just in answering questions — it's in taking action. Our agents have access to a registry of tools they can invoke mid-conversation. When you ask a question that requires live data, document lookups, or write operations, the agent autonomously decides which tools to call, in what order, and how to synthesize the results.
This is function calling at the LLM layer. The model doesn't execute code — it emits structured JSON describing which tool to call and with what parameters. Our server-side tool registry handles execution and returns results back into the conversation context.
Pulls real-time futures quotes from the Polygon API. Ask "What's gold trading at?" and the agent
calls execute(symbol: "GC") to fetch the latest
price, change, and volume. No stale training data — live market data streamed into the conversation.
Performs semantic search across your entire document library. Uses the same embedding + cosine similarity pipeline as RAG, but invoked on-demand by the agent. Useful when the initial RAG context isn't enough and the agent needs to dig deeper.
Generates summaries, outlines, and concept maps from any source document. Results are cached so the first analysis is LLM-powered, and subsequent requests are instant. Ask "Summarize that Goldman Sachs copper report" and the agent handles the rest.
Agents can read and write to your research stories.
Mid-conversation, you can say "Add this analysis to my Nat Gas Winter Thesis story" and the agent
creates a timestamped journal entry. ListStoriesTool lets it browse your
full story inventory with associated tickers and recent entries.
Resolves @mentions in your messages. Reference a ticker, story, position,
or company by name and the tool injects the full context — current prices, thesis details,
P&L, fundamentals — directly into the conversation. It's how agents "see" the entities
you're talking about.
Right Model for the Right Job
We don't believe in a one-model-fits-all approach. Different tasks have fundamentally different cost-latency-quality tradeoffs, and we route accordingly. Here's how we allocate model capacity:
| Task | Model | Why |
|---|---|---|
| Interactive Chat | gemini-2.5-flash | Fast streaming, good tool use, low cost per token |
| CFTC & EIA Report Generation | gemini-2.5-pro | Complex tabular reasoning, long-context analysis |
| Story Ticker Enrichment | gpt-4o-mini | High throughput, good at structured reasoning |
| Metadata Extraction | gemini-2.5-flash | Schema-constrained output, fast, cost-efficient |
| Embeddings | text-embedding-3-small | 1536 dims, best cost/quality for semantic search |
| Graph Visualization | gemini-2.5-flash | Creative structured output, Mermaid syntax generation |
We use RubyLLM as our abstraction layer across providers. This gives us a unified interface for chat, embeddings, and tool use across OpenAI and Google Gemini, with the ability to swap or add providers without changing application code. Model metadata — context windows, capabilities, pricing — is tracked in our database so we can route dynamically.
Government Data to Trading Intelligence in Minutes
The CFTC Commitments of Traders report and the EIA Natural Gas Storage report are two of the most watched data releases in commodity markets. We've built automated pipelines that fetch, parse, embed, and analyze these reports the moment they drop.
flowchart TB
subgraph Schedule["Scheduled Triggers"]
A["CFTC: Fridays @ 2:30 PM ET"]
B["EIA NG: Thursdays @ 10:30 AM ET"]
end
subgraph Fetch["Fetch & Transform"]
A --> C["FetchCFTCReportsJob"]
B --> D["FetchEIANaturalGas
StorageReportsJob"]
C --> E["Create FeedRun
state: processing"]
D --> E
E --> F["fetch_data → raw"]
F --> G["transform_data → normalized"]
G --> H["validate_schema"]
H --> I["Store result_data
as JSONB"]
end
subgraph Embed["Embed & Index"]
I --> J["GenerateFeedEmbeddingsJob"]
J --> K["Feed-Specific
Chunking Strategy"]
K --> L["Concurrent Embedding
10 fibers"]
L --> M[("FeedChunk records
in pgvector")]
end
subgraph Analyze["AI Report Generation"]
I --> N["GenerateCFTCReport
or GenerateEIAReport"]
N --> O["Gemini 2.5 Pro
with specialized prompt"]
O --> P["Markdown Report
with analysis"]
P --> Q["Published to
Reports feed"]
end
style Schedule fill:#1e1b4b,stroke:#818cf8,color:#ffffff
style Fetch fill:#312e81,stroke:#818cf8,color:#ffffff
style Embed fill:#1e1b4b,stroke:#818cf8,color:#ffffff
style Analyze fill:#312e81,stroke:#818cf8,color:#ffffff
Automated pipeline: scheduled fetch, schema validation, concurrent embedding, and AI report generation.
By the time you sit down Monday morning, last Friday's CFTC data has already been:
No manual parsing. No spreadsheet wrangling. The intelligence is there when you need it.
Beyond Vector Search: Structured Relationships
Vector embeddings are powerful for semantic similarity, but they don't capture relationships. Knowing that a document chunk about "copper mine expansions" is semantically close to "Freeport-McMoRan" is useful. Knowing that Freeport-McMoRan operates the Grasberg mine, which produces copper and gold, and that the CEO discussed capex guidance on the Q3 call — that's a knowledge graph.
We use Apache AGE (A Graph Extension) running on PostgreSQL to store entity-relationship triples alongside our vector store. This means a single query can combine:
graph LR
A(("Freeport-
McMoRan")) -->|OPERATES| B(("Grasberg
Mine"))
B -->|PRODUCES| C(("Copper"))
B -->|PRODUCES| D(("Gold"))
A -->|CEO| E(("Kathleen
Quirk"))
E -->|DISCUSSED| F(("Q3 2025
Earnings Call"))
F -->|MENTIONS| G(("Capex
Guidance"))
F -->|MENTIONS| C
C -->|TRACKED_BY| H(("CFTC COT
Report"))
H -->|SHOWS| I(("Spec Net Long
+42k contracts"))
style A fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style B fill:#312e81,stroke:#818cf8,color:#ffffff
style C fill:#b45309,stroke:#fbbf24,color:#ffffff
style D fill:#a16207,stroke:#fbbf24,color:#ffffff
style E fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style F fill:#312e81,stroke:#818cf8,color:#ffffff
style G fill:#312e81,stroke:#818cf8,color:#ffffff
style H fill:#1e3a5f,stroke:#60a5fa,color:#ffffff
style I fill:#1e3a5f,stroke:#60a5fa,color:#ffffff
Example knowledge graph: entities, relationships, and data sources connected across your research.
When our agents process documents, they extract structured entities into the graph. Typical node labels include:
PERSON
Analysts, executives, officials
ORGANIZATION
Companies, exchanges, regulators
CONCEPT
Contango, backwardation, carry
TOPIC
Supply disruption, weather, geopolitics
Streaming Architecture: No Waiting for Full Responses
When you send a message, you don't wait 30 seconds for a complete response. Tokens start appearing in your chat window within milliseconds of the first LLM output. Here's how:
sequenceDiagram
actor User
participant Controller as Rails Controller
participant Sidekiq as Sidekiq Job
participant LLM as LLM Provider
participant Turbo as Turbo Streams
User->>Controller: POST /chat_messages
Controller->>Controller: Create ChatMessage record
Controller->>Sidekiq: ChatStreamJob.perform_later(chat_id)
Controller-->>User: Turbo Stream: append message shell
Sidekiq->>Sidekiq: Register tools from ToolRegistry
Sidekiq->>LLM: chat.complete (streaming)
loop Token by token
LLM-->>Sidekiq: chunk
Sidekiq->>Turbo: broadcast_append_chunk
Turbo-->>User: Real-time DOM update
end
Note over Sidekiq,LLM: Tool calls happen mid-stream:
Agent pauses, calls tool, resumes generation
Sidekiq->>Sidekiq: Save final assistant message
Sidekiq
Chat streaming runs in dedicated background workers. Your browser isn't blocking on LLM inference.
Turbo Streams
Hotwire's Turbo Streams push DOM updates over WebSocket. Each token chunk triggers a targeted HTML append.
Zero JS Bloat
No React. No client-side state management. Server-rendered HTML streamed in real-time via Rails conventions.
Every Entity Is a First-Class Citizen
In Arc Research, tickers, stories, positions, companies, and users are all "mentionable" entities.
When you type @GC (gold futures) or @My Copper Thesis
in a chat message, the ChatContextInjectionService resolves those references
and injects the full entity context into the conversation.
This means the agent isn't just reading text — it knows what you're talking about. If you mention a story, the agent sees the story's title, associated tickers, and the latest journal entry. If you mention a ticker, it gets the current quote, your position (if any), and relevant research.
flowchart LR
A["User message:
'How does @GC positioning
look vs my @Gold Thesis?'"] --> B["Context Injection
Service"]
B --> C["Resolve @GC
→ Ticker context"]
B --> D["Resolve @Gold Thesis
→ Story context"]
C --> E["GC1! $2,340.50
Vol: 234k | OI: 512k"]
D --> F["Story: Gold Thesis
Tickers: GC, GLD, NEM
Last entry: Feb 12"]
E --> G["Enriched Prompt
to LLM Agent"]
F --> G
G --> H["Agent response with
full entity awareness"]
style A fill:#312e81,stroke:#818cf8,color:#ffffff
style B fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style C fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style D fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style E fill:#1e3a5f,stroke:#60a5fa,color:#ffffff
style F fill:#1e3a5f,stroke:#60a5fa,color:#ffffff
style G fill:#4338ca,stroke:#a5b4fc,color:#ffffff
style H fill:#312e81,stroke:#818cf8,color:#ffffff
The Infrastructure Stack
We built this on battle-tested technology. No experimental frameworks. No hype-driven decisions. Every component was chosen because it's the right tool for the job in a latency-sensitive, data-intensive research application.
Why This Architecture Matters for Trading
You could use a general-purpose AI chatbot. You could copy-paste CFTC data into a spreadsheet. You could manually search through PDFs and broker notes. Traders have been doing these things for years. Here's why this is different:
Every document you add, every feed that runs, every journal entry you write makes the system smarter for your specific research workflow. It's not starting from zero each conversation — it's building on everything you've fed it. Think of it as compound interest for your research graph.
When you ask about copper positioning, the agent can simultaneously reference your uploaded Goldman Sachs research PDF, the latest CFTC feed data, your story about Freeport-McMoRan, and the live CL quote — all in one response. No human analyst can context-switch across that many sources that fast.
CFTC data hits at 2:30 PM Friday. EIA storage at 10:30 AM Thursday. Our background jobs are already running when the data drops — fetching, validating, embedding, and generating analysis. By the time you check in, the work is done.
Your research is your edge. Every vector embedding is user-scoped. Your documents, your stories, your positions never leak into another user's context. The knowledge graph is yours alone.
With Arc Research Starter, the same knowledge graph and market tools are available over MCP in Cursor and Claude. Generate a token, connect once, and keep researching where you already work. See MCP setup.
When a detector fires, Archie can propose a paper trade, send it through review, and leave a run you can audit. Hard stops stay in code. Live tickets stay gated. Open Archie Desk.
Start building your personal knowledge graph. Upload a document, chat with an agent, connect MCP from Cursor, or open Archie Desk and read the next paper decision.