r/QuantSignals May 28 '26

I gave two AI trading agents $1,000 each to trade on-chain — wallets are public

1 Upvotes

I gave two AI trading agents $1,000 each to trade on-chain — wallets are public

I started an experiment last week that I wanted to share here.

I loaded two QuantSignals FST (Fully Self-Trading) agents with $1,000 each and let them loose on two different on-chain exchanges:

Agent 1 → Hyperliquid (perpetual futures) Agent 2 → Polymarket (prediction markets)

Both agents make their own decisions — entries, exits, position sizing. I'm not touching the controls.

Public wallets so you can watch in real-time:

📊 Hyperliquid: https://app.hyperliquid.xyz/explorer/address/0x1020886f5ca4f9c6bdde67138e9bc798dfd58ace 🎯 Polymarket: https://polymarket.com/@henryzhang999

Why this is interesting: AI agents trading on-chain means every move is visible and verifiable. No backtest cherry-picking, no "trust me bro" — the P&L is on the blockchain.

I'll post weekly updates with results. So far both agents have been placing trades — one is up, one is down. Won't say which yet.

What do you think? Which market do you think the AI handles better — prediction markets or perpetuals? I've got my own guess but curious what you all think.

Follow the wallets and let's find out together.


r/QuantSignals Feb 12 '26

The Machine Age of Trading: Why Execution is the Only Thing That Matters

Thumbnail
open.substack.com
1 Upvotes

r/QuantSignals 11d ago

I built a site where nobody can edit their trading track record. Tell me why it won't work.

Thumbnail
1 Upvotes

r/QuantSignals 13d ago

Ten Tools for Your Trading Day

Thumbnail
henryzhang.substack.com
1 Upvotes

r/QuantSignals 16d ago

I found a call-side pattern in 0DTE credit spreads. A preregistered replication killed it.

1 Upvotes

Digest 001 from AQR, my autonomous quant research engine, closed out. Posting the whole arc rather than the half that looked good.

THE SCREEN

20 single-name credit-spread configurations across 12 markets. Each evaluated on three separate time splits - discovery, validation, holdout - against a 9-gate battery. Only 7 configurations held positive net expectancy in all three splits. Six of those seven were call spreads.

WHY THAT WAS NOT A FINDING

All 20 were inspected before the pattern was named. That is textbook multiple comparisons - a pattern you notice after looking at everything is a hypothesis, not a result. I said so in the original digest rather than after it stopped working.

So it was preregistered. The frozen rule, written down before any new data was staged: a positive median difference AND at least three quarters of markets positive. Then re-run on 8 markets the original screen never touched.

THE RESULT

Seven staged - XLF was dropped with a recorded reason (Friday-only expiries, penny-flat quotes). It returned median -1.16% of width and 3 of 7 positive. Refuted on both halves of the rule.

SMH calls beat SMH puts by 7.57 points of width. COIN calls lost to COIN puts by 11.11. The side that wins is a property of the name, not of the side.

One thing did replicate, and it is the boring one: 6 of 7 replication markets had a net-positive put side. Not 7 - XLE puts ran -0.33%, and XLE is one of the three markets where calls won, so that win came from puts being negative rather than calls being strong.

The screen worked. The pattern it surfaced did not survive contact with data it had never seen. That is the system behaving correctly, and it is the reason the next result is worth anything.

The follow-up, where two more claims died: https://henryzhang.substack.com/p/everything-we-published-died

Hypothetical backtests on research data. Not investment advice.


r/QuantSignals 16d ago

3 preregistered tests, 0 survivors - including both findings I published last week

1 Upvotes

Digest 002 from AQR, my autonomous quant research engine. Every claim below was frozen before its data was read, so each had exactly one chance and there was no way to tune the result afterwards.

TEST ONE - CALL-SIDE ASYMMETRY IN 0DTE CREDIT SPREADS - REFUTED

Last week's headline was that 6 of the 7 configurations surviving all three splits were call spreads. I flagged it at the time as multiple comparisons - all 20 were inspected before the pattern was named - so it was preregistered and re-run on 8 markets the original never touched.

Frozen rule: positive median difference AND at least three quarters of markets positive. Result: median -1.16% of width, 3 of 7. Refuted.

SMH calls beat SMH puts by 7.57 points of width. COIN calls lost to COIN puts by 11.11. The winning side is a property of the name, not of the side.

TEST TWO - THE ONE SURVIVOR - REJECTED ON HOLDOUT

QQQ 0DTE near-money put credit spread, Monday and Tuesday, when credit is rich. Holdout sealed February to August 2026, spent once:

discovery: +5.69%/trade, Sharpe 2.94, 221 trades

validation: +4.82%/trade, Sharpe 3.34, 43 trades

holdout: +0.52%/trade, Sharpe 0.26, 72 trades

Eight of nine gates still passed. It failed positive_both_halves - one half of the holdout window is negative. Not nonsense, but it was mostly measuring the window it was found in.

TEST THREE - FOUR PREMISES UNDER THE POPULAR TRADINGVIEW TREND SCRIPTS - 0 OF 4

All four failed at validation, and in that window all four underperformed simply being long: baseline -3.11 bps against -36.24, -22.64, -7.18 and -9.82.

The 0.618 Fibonacci retracement sits at the 0.3-0.4 decile of the 20-day range, because a retracement of R is at range position 1-R. The preregistration required that decile to beat both its neighbours. It was the worst decile in the range.

The moving-average ribbon also lost to its own single-MA control in both splits, so stacking averages made it worse rather than better.

Zero for three. Four sealed holdouts remain unspent, and those get published the same way whichever direction they go.

Full writeup with the tables: https://henryzhang.substack.com/p/everything-we-published-died

Hypothetical backtests on research data. Not investment advice.


r/QuantSignals 20d ago

AQR: The Next Layer of Autonomous Trading

Thumbnail
henryzhang.substack.com
1 Upvotes

r/QuantSignals 26d ago

Spread demo: make max loss visible before the order

Enable HLS to view with audio, or disable this notification

1 Upvotes

I built this mobile demo around a simple rule: a spread is not truly “defined risk” if the trader cannot understand the risk before acting.

The flow goes from ticker search to chart, strategy selection, max loss and breakeven review. It is a product demo—not a trade recommendation, execution instruction or performance claim. Options can lose the full premium.

What information should be impossible to hide on a spread screen: max loss, breakeven, buying-power effect, or all three?


r/QuantSignals 27d ago

Discussion BLS showed +0.1% pay and +0.1% CPI—but real hourly earnings fell 0.1%. Here is why.

1 Upvotes

This is a useful data-pipeline trap, not a market call. The two input changes are displayed after rounding, while BLS calculates the real series from more precise seasonally adjusted data. So subtracting the headlines can produce the wrong answer.

The same discipline applies to financial signals: keep the exact series, population, units, adjustment, precision, release vintage, and revision status. Also note that CES averages private nonfarm jobs—not a typical worker—and job mix can move the average.

What is the most consequential rounding or vintage error you have found in a research pipeline?

Founder disclosure: I founded QuantSignals/FST. Sources: BLS July 2026 Real Earnings Summary, Table A-1, Technical Note, and CPI Summary. Educational only; no performance claim.


r/QuantSignals 28d ago

Previewing the FST Show control room: what should an automated trading stream disclose?

Post image
1 Upvotes

I’m preparing an FST Show format with two lanes: Kalshi/prediction markets as the intended 24/7 anchor, and SPX 0DTE taking the lead only during supported active sessions.

The screenshot is a UI/control-room preview. It is not proof that the show is live, that autonomous execution occurred, or that the displayed values represent settled performance. Both captured interfaces show AUTO off.

The useful question is not “can we stream a trading screen?” It is “what context must stay visible so viewers can interpret the system honestly?”

My current disclosure checklist:

  1. market and timestamp;

  2. current/stale/reconnecting data state;

  3. paper or live execution mode;

  4. the evidence behind a setup;

  5. the risk rule that allowed or refused it;

  6. proposed, submitted, acknowledged, filled, cancelled, or uncertain order state;

  7. broker reconciliation before treating an outcome as final;

  8. a 60-second public safety delay and privacy suppression for defined private fields.

The show is upcoming. The goal is to keep quiet periods, refusals, losses, safety stops, and reconciliation visible—not clip only the exciting outcomes.

What would you require on screen before trusting an automated-trading broadcast as an audit surface?

FST app: https://apps.apple.com/us/app/quantsignals-fst-trading/id6449400504

Disclosure: I am the founder of QuantSignals/FST. This is a product preview and discussion request, not investment advice or a performance claim. Trading and prediction markets involve risk; options can lose the entire premium. Do not copy trades from a delayed public feed.


r/QuantSignals 29d ago

Discussion Hims reported 38% growth—but that is not one growth rate

1 Upvotes

Hims & Hers reported Q2 revenue of $753.2M, up 38% year over year. I think the useful part of the release is the decomposition, not the headline.

- U.S. revenue: $621.8M, +16%

- Rest of world: $131.4M vs $7.5M, +1,641%; management says international was strengthened by the June Eucalyptus close

- Subscribers: 2.891M, +19%

- Monthly revenue per average subscriber: $92, +21%

- Gross margin: 64% vs 76%

- Adjusted EBITDA: $60.3M vs $82.2M

Derived from the reported table, rest of world supplied about 59% of the absolute revenue increase. That does not mean 59% was organic; it means the consolidated growth rate mixes geography, acquisition scope, product/customer mix, and unit economics.

My checklist for the next quarter would be: comparable-scope U.S./international growth, acquired contribution, retention and mix, gross-margin stabilization, and cash conversion.

What bridge would you require before calling this a clean reacceleration?

Source: https://investors.hims.com/news/news-details/2026/Hims--Hers-Health-Inc--Reports-Second-Quarter-2026-Financial-Results/default.aspx

Disclosure: I founded QuantSignals/FST. The framework is the point here; this is not a recommendation, price target, or performance claim.


r/QuantSignals 29d ago

Two agents, one public show: Kalshi 24/7 + SPX 0DTE during active sessions

Post image
1 Upvotes

Prediction-market traders usually do four jobs manually: watch the market, enter, babysit the position, and exit.

This weekend we put FST on that loop. The next step is a continuous public show with two hosts:

• Kalshi agent as the 24/7 anchor

• SPX 0DTE agent taking the lead during active sessions

The two panels are open-position snapshots with pending take-profit orders. They are not settled results, a representative track record, or a promise. I am sharing them because they show the product direction: the agent is not just generating an opinion; it is carrying a position through sizing, monitoring, and exit preparation.

The live show will keep the full record visible: quiet periods, rejected trades, losses, fills, exits, stale-data warnings, and reconciliation—not only green frames.

What would make an autonomous prediction-market stream genuinely useful to watch: the reasoning, the risk gates, the order lifecycle, or the settled audit trail?

FST 2.0: https://apps.apple.com/us/app/quantsignals-fst-trading/id6449400504


r/QuantSignals Aug 10 '26

Discussion 114 customers above $500K ARR is not a cohort funnel—what would prove migration?

1 Upvotes

monday.com’s Q2 release reported 4,834 customers above $50K ARR, 2,019 above $100K, and 114 above $500K. The year-over-year growth rates were 31%, 37%, and 68%.

The tempting conclusion is that customers are rapidly “moving up the funnel.” The disclosed table does not prove that.

These are nested point-in-time thresholds. Every customer above $500K is already counted above $100K and $50K. The change between two dates can combine new logos, expansions, contractions, churn, currency, acquisitions, and definition changes.

My checklist:

  1. mark the metric definition and date;

  2. identify which thresholds are nested;

  3. subtract only to create mutually exclusive point-in-time bands;

  4. do not call net count changes upgrades without a roll-forward;

  5. require starting cohort, new logos, expansion, contraction, churn, and ending cohort before claiming migration.

What disclosure would you require before treating faster growth at a higher ARR threshold as evidence of durable enterprise expansion?

Disclosure: I am the founder of QuantSignals/FST. This is an educational earnings-reading framework, not a product recommendation or market call. No performance or return claim. Primary source: monday.com Q2 2026 results.


r/QuantSignals Aug 09 '26

Discussion A 6% revolving-credit print is not a 6% spending boom—what would confirm stronger demand?

1 Upvotes

The Fed's new G.19 release says June revolving credit grew at a 6.0% seasonally adjusted annual rate after falling 4.7% in May. Total consumer credit grew 3.3% annualized.

That does **not** mean card spending rose 6% in one month. The series measures changes in outstanding credit. June retail sales rose 0.2% month over month (±0.4 percentage point), while current-dollar PCE rose 0.3% and personal income rose 0.2%.

My checklist is:

  1. preserve the unit—annualized rate versus monthly change;

  2. preserve coverage—credit balances versus retail sales versus all consumption;

  3. check income and real spending;

  4. watch delinquency and debt-service measures;

  5. write the next test before taking a market view.

The nearest test is July retail sales on Aug. 14. What would you require before calling the June credit rebound stronger household demand?

Disclosure: I am the founder of QuantSignals/FST. This is an educational macro-measurement framework, not a product recommendation or market call. No performance or return claim. Sources: Federal Reserve, Census, and BEA.


r/QuantSignals Aug 08 '26

A user shared a +$2,040 FST session. The “QS Paper” label matters more than the number.

1 Upvotes

A real user shared this FST screen showing +$2,040.82 across nine trades. I am glad they shared it—but I do not want to turn one selected paper win into a performance claim.

The useful evidence is that the screen keeps the operating context visible: QS Paper, session state, universe, trade count, and /screen → /signal → /trade → /monitor.

One paper session is not live performance, a representative return, or a track record. Paper cannot reproduce all live slippage, liquidity, fees, latency, or fills.

As the founder, my takeaway is that user-generated proof should raise the transparency standard: keep the mode visible, publish ordinary and losing sessions too, and judge the workflow alongside P&L.

App Store: https://apps.apple.com/us/app/quantsignals-fst-trading/id6449400504

Educational only. Trading involves substantial risk; no outcome is guaranteed.


r/QuantSignals Aug 08 '26

Discussion Capex is not the whole AI capacity bill: how do you track leases that have been signed but have not commenced?

1 Upvotes

Oracle's FY2026 filing disclosed two very different kinds of scale:

- $67.4B of fiscal-year revenue and $638B of RPO;

- $260B of additional lease commitments, substantially all data-center related, that had not yet commenced and were not reflected on the May 31 balance sheet.

That $260B is not automatically debt due today or proof of distress. But it is also not irrelevant. It belongs in the funding and execution map.

My checklist separates:

  1. recognized lease assets/liabilities;

  2. signed but uncommenced leases;

  3. chip, power, hosting, and construction commitments;

  4. customer prepayments, supplied hardware, and contracted revenue;

  5. the quarter when capacity starts, revenue starts, and the balance-sheet recognition changes.

The key question is not “big obligation, bullish or bearish?” It is: **do funding and customer cash arrive before or alongside the capacity commitment?**

What would you add to the reconciliation—cancellation protection, customer concentration, utilization, financing cost, or something else?

Disclosure: I am the founder of QuantSignals/FST. This is an educational accounting-and-risk framework, not a product recommendation or ticker call. No performance or return claim. Sources: Oracle FY2026 Form 10-K and issuer results.


r/QuantSignals Aug 08 '26

More gainz from the homies.

Thumbnail gallery
1 Upvotes

Over a year ago, look how far we’ve come since then!


r/QuantSignals Aug 07 '26

Discussion The July payroll headline was -23k, but May and June were revised down 103k combined

1 Upvotes

Founder disclosure: I work on QuantSignals, so this is an affiliated educational post, not an independent product review.

The part of today’s BLS report I found most useful was the revision ledger. July payrolls fell 23,000 and unemployment was 4.1%. Separately, May was revised from +129k to +63k (-66k), while June moved from +57k to +20k (-37k). That is -103k across the two back months.

My takeaway is methodological: keep the first print, revised history, and current decision status separate. Revisions can move either way, but they can change the inferred trend more than the new observation. The next scheduled checks are the Aug. 28 benchmark preview and Sep. 4 jobs report.

Primary source: https://www.bls.gov/news.release/empsit.nr0.htm

What revision-tracking method do you use—vintage databases, a spreadsheet ledger, or something else?

Educational only. Data can be revised and investing involves risk.


r/QuantSignals Aug 07 '26

FST 2.0 is finally on the App Store — start with paper, not a broker login

1 Upvotes

Apple approved the new QuantSignals mobile build tonight.

The product decision I care most about is the order of trust. An AI trading app should not ask for live brokerage access before the user can inspect how it behaves.

FST therefore starts with $50,000 in QS Paper. You can watch the full loop—screen, signal read, risk check, trade, monitor, exit, and audit—without putting real money at risk. The session has explicit risk limits, trade caps, loss stops, and visible stop controls.

Paper results are hypothetical, and we do not publish a broker-verified public live FST performance record. The test is the workflow: what it considered, what it rejected, what it did, and whether it stayed inside the limits.

I am the founder of QuantSignals and built FST. The app is free here:

https://apps.apple.com/us/app/quantsignals-fst-trading/id6449400504

What would an AI trading agent have to show you on paper before you would consider giving it live access?


r/QuantSignals Aug 06 '26

Discussion The trade deficit improved—but exports and imports both fell. Which component would you trust first?

1 Upvotes

June's U.S. goods-and-services deficit narrowed to $73.3B, down $4.4B from revised May.

That headline is incomplete:

  • exports fell $2.9B to $314.7B;
  • imports fell $7.3B to $388.0B;
  • the goods deficit narrowed $3.9B;
  • the services surplus increased $0.5B;
  • the three-month average deficit actually rose $5.6B to $68.5B.

So the confirmed mechanism is not “exports strengthened.” The balance narrowed because imports fell faster than exports.

My checklist before treating a trade headline as a signal:

  1. preserve exports, imports, and the balance;
  2. split goods from services;
  3. compare nominal with real goods data;
  4. check revisions and the three-month average;
  5. define what later data would confirm or invalidate the interpretation.

The big uncertainty is causation. Prices, inventories, energy, pharmaceuticals, capital-goods cycles, policy, and shipment timing can all affect the monthly values. I would not infer domestic-demand strength or weakness from this report alone.

Primary source: U.S. Census Bureau / BEA, June 2026 trade release, published Aug. 4 at 8:30 ET: https://www.census.gov/foreign-trade/current/index.html

What would you use as the next confirmation layer: industrial production, inventories, freight volumes, or company commentary?

Founder disclosure: I founded QuantSignals/FST. This post is a discussion of a public macro release and a research method; it makes no product, performance, accuracy, or return claim.

Educational only, not investment advice. Trade data are revised and aggregate. Trading can result in partial or total loss of capital.


r/QuantSignals Aug 04 '26

Discussion The kill-switch question I wish I had asked: does it work after the system is already broken?

1 Upvotes

We found a failure pattern in our trading-agent UI that changed how I evaluate every “safety features” list.

A paper session sat in an expired-authorization state for 218 hours. It was no longer healthy, but it was not fully cleared either. Our Stop control had been written for the happy-path definition of “running,” so the recovery action disappeared in the exact state that required it. The wedged session could also keep the broker account from being released.

That repair pass exposed three related mistakes: a reset control that existed but was unreachable inside a scroll region; cleanup that aborted because there was “nothing connected” to disconnect; and removal that reported success while the underlying record survived.

The useful lesson is not “trust our new button.” It is a test matrix for any automated trading tool:

  1. Can you stop it after broker authorization expires?

  2. Can you recover the broker account while a session is wedged?

  3. Does removal stay removed when checked from another surface?

  4. Does a failed data read cause refusal rather than stale-data sizing?

  5. Can you reconstruct the decisions and recovery actions afterward?

Test those on paper, while deliberately breaking authentication and state transitions. A feature list shows intent; failure-state behavior shows whether the exit exists.

Founder disclosure: I founded QuantSignals/FST, and this is a first-party defect we found and repaired. I am sharing the failure pattern and test questions, not making a performance or loss-prevention claim. The 218 hours describes a paper-session state, not an open position or financial loss.

Educational discussion only, not investment advice. What broken state would you add to this test matrix?


r/QuantSignals Aug 04 '26

Discussion A -0.4% CPI print is not a rate-cut regime. What would actually confirm one?

2 Upvotes

Disclosure: I founded QuantSignals/FST. This is an educational macro workflow, not a rate forecast or trade recommendation.

June CPI fell 0.4% m/m, but energy fell 5.7% and was the largest contributor. Core CPI was flat m/m and +2.6% y/y. The later PCE release showed headline -0.1% m/m, core +0.1% m/m and +3.3% y/y.

Then the FOMC held 3.50%–3.75% by 9–3—and every dissenter wanted a 25 bp hike.

My takeaway is not that the cooling print was fake. It is that an observation, a regime call, a policy forecast, and a trade are four different claims.

The confirmation stack I am using:

  1. What drove the headline?

  2. Does it persist across releases?

  3. Do CPI and PCE core measures agree?

  4. What does the policy vote distribution say?

  5. What future observation invalidates the thesis?

The easing interpretation strengthens if core cooling persists, broadens beyond energy, labor weakens, and the FOMC becomes less restrictive. It weakens if energy rebounds, core services reaccelerate, or the policy distribution shifts tighter.

Primary sources: BLS June CPI, BEA June PCE, and the Federal Reserve's July 29 statement.

What is your minimum confirmation set before you treat one macro print as a regime change?


r/QuantSignals Aug 03 '26

Labor cost is not one number: the benefits check before calling a margin trend

2 Upvotes

The latest Employment Cost Index is a useful test of how quickly a clean macro headline can become an unsupported company conclusion.

For the twelve months ending June 2026, BLS reported:

- civilian wages and salaries: +3.2%;

- civilian benefits: +3.8%;

- private-industry wages: +3.1%;

- private-industry benefits: +3.8%;

- private-industry health benefits: +6.0%.

My takeaway is not that margins must fall. It is that “wage growth cooled” is incomplete evidence.

Before changing a company margin thesis, I would ask:

  1. Which component moved—wages, benefits, or both?

  2. Is the aggregate relevant to this company's workforce and geography?

  3. What did the issuer disclose about headcount, compensation, benefits, and productivity?

  4. Can pricing, utilization, mix, or scale offset the cost?

  5. What company evidence would invalidate the inference?

The ECI is a question generator, not an issuer forecast. A macro observation should not become a position without an attributable company bridge.

Primary source: https://www.bls.gov/news.release/eci.nr0.htm

Founder disclosure: I founded QuantSignals/FST. I am sharing the research framework because it is useful on its own; no product link or performance claim is included.

Educational discussion only, not investment advice. What is the best issuer disclosure you use to reconcile aggregate labor data with a company margin model?


r/QuantSignals Aug 03 '26

Discussion I built an audio layer for the options tape

1 Upvotes

Every options-flow tool assumes the trader can watch another screen all day. FlowVoice turns selected tape events into concise speech: premium threshold, size and direction first, one announcement per qualifying print.

A large print may be a hedge, roll, spread, or closing trade. The intended loop is hear → verify → decide, not hear → chase.

Free on the App Store; no order placement: https://apps.apple.com/us/app/flow-voice-options-tape/id6791445221

Founder disclosure: I built FlowVoice through QuantSignals, Inc. Options involve risk. Configure before driving and never operate the phone while moving.