IPO Offering Structure: A Methodology for Reading Pricing, Fees, and Underwriter Composition From the Prospectus

An FMP Signals Lab Research Methodology

Version 1.0. Effective August 2026.

What This Methodology Measures

Most IPO coverage is built around two numbers: the valuation and the first-day move. A company goes public at a headline valuation, the stock rises or falls on its first day, and that move becomes the story. Both numbers are real. Neither describes how the offering was constructed.

Offering structure is a different question. It asks what the terms of the offering disclose about the deal itself: what the underwriters charged to bring it to market, how much of each share's price reached the company, how much of the business was sold to the public, whether existing shareholders sold alongside the company, and which banks underwrote the deal. These are structural facts, fixed before the stock trades, and every one of them is disclosed in the final prospectus.

This methodology standardizes those terms so that offerings of different sizes and structures can be read on the same basis. It is deliberately descriptive. It reports what the documents state. It does not assign a score, rank deals by quality, or predict how a stock will trade. A low underwriting spread may reflect the scale of the offering rather than the quality of its execution. A small offered-share ratio may reflect a deliberate structure rather than a defect. Those interpretations belong to analysis built on this methodology, not to the methodology itself.

This page is the reference for that analysis. It defines the universe, the measures, the retrieval process, and the validation the process requires. Case studies published separately apply it to specific offerings and link back here rather than restating it.

How the Methodology Is Run

The framework is run as a recurring quarterly process. Each cycle covers the IPOs priced during one calendar quarter, and the same universe rule, measures, retrieval process, and validation requirements are applied to every offering in that quarter. This page holds the definitions. It does not hold current-quarter findings, which are published separately.

A quarterly universe is treated as complete on a fixed schedule rather than when the data appears sufficient. The universe closes forty-five calendar days after the last day of the quarter. That interval is set by the retrieval process itself: a final prospectus is filed at or shortly after pricing under Rule 424, and the methodology treats a filing as an offering's IPO prospectus only when it is dated within forty-five days of the pricing date, so waiting a full forty-five days past quarter end gives every offering in the quarter the same opportunity to be matched. Offerings priced in the quarter whose prospectus has not appeared by that date are carried in the universe as unmatched rather than excluded, because their absence is a fact about the quarter and dropping them would understate it.

Every study states the snapshot date on which its data was retrieved, and all figures in that study are as of that date. Filings amended after the snapshot are not retrofitted into a published study. If a correction is material, it is issued as a dated revision to that study rather than applied silently, so that a reader returning to a study sees the same figures the study reported.

The Limits of Headline IPO Coverage

First-Day Performance Measures Market Reaction

The first-day move measures demand on a single day, shaped by allocation, marketing, and market conditions at that moment. It is a reading of reaction, not of construction. A deal that opens well above its offering price transferred value from the company to the buyers who received allocations, which may reflect cautious pricing rather than a well-run offering. A deal that opens close to its price may have been priced efficiently. The move describes the first afternoon of trading. It says nothing about what the company paid to go public, how much of itself it sold, or who underwrote the deal.

Valuation has the same limitation in a different form. It is a fact about scale. Two companies can list at similar valuations while one sells a substantial share of itself at a low spread with no secondary component and the other sells a small fraction at a high spread with existing shareholders participating. The headline treats those as equivalent. They are structurally different offerings.

Structural Information Lives in the Prospectus

The terms that describe the offering are set at pricing and disclosed in the final prospectus, which is filed at or shortly after pricing under Rule 424. The cover page carries the offering price, the underwriting discount, and the proceeds before expenses. The body carries the share counts by class, the split between company and selling-shareholder shares, the over-allotment option, and the full underwriting syndicate with each bank's role.

That document is the authoritative source for everything in this methodology. Structured data covering the same fields is convenient where it exists, and this methodology uses it as a cross-check, but the filing is what the analysis rests on.

Standardized Measures Make IPOs More Comparable

The reason to standardize is comparability. A gross spread, a retained fraction, an offered-share ratio, and a syndicate profile can be computed the same way for every deal where the underlying disclosure is available, then read side by side. That turns a set of individual narratives into a consistent description that holds up across offerings of very different sizes.

Comparability has boundaries, and this methodology states them rather than assuming them away. Ratios computed against different share-class denominators are not comparable to each other. Figures extracted from a document without validation are not comparable to figures confirmed against the filing. The sections on engineering and on where the methodology breaks down set out where those boundaries fall.

The Data Behind Offering-Structure Analysis

Structured IPO Economics

The cover-page economics, offering price per share, underwriting discounts and commissions, proceeds before expenses, and total offering amount, are available as structured fields. Where a record exists for a deal, the gross spread and the retained fraction follow directly, with no document handling required. These are the cheapest measures in the methodology and the easiest to compute consistently.

Records are not present for every offering, and the endpoint that carries them also carries prospectus filings for later offerings by the same companies. Both conditions are handled explicitly rather than assumed away, and the next subsections describe how.

Prospectus-Derived Signals

Three of the measures this methodology reports are not available as structured fields at all: the underwriting syndicate, the share counts that an offered-share ratio requires, and the split between company shares and selling-shareholder shares. All three are disclosed in the prospectus, in prose and in tables rather than in a single machine-readable field.

Reaching them means retrieving the document and reading it. This is not a shortcoming of any particular data source. The disclosures themselves are textual, and the filing is where they live.

A Staged Extraction Process

The process runs in a fixed order. The offering calendar defines the universe. The filings record locates each deal's final prospectus. Structured economics attach to the deals that have a record. The document is then retrieved and scanned, producing candidate facts. Those candidates are validated against the filing before anything is reported.

The order matters because each stage narrows what the next one has to handle, and because the validation step sits at the end rather than being skipped. A figure that has not been checked against its context in the filing is a candidate, not a measurement.

Missing Data and Validation

Three conditions recur often enough to be treated as normal rather than exceptional. A deal may appear in the calendar with no structured economics record, in which case the economics are read from the filing. A document scan may return several plausible values for the same quantity, in which case the correct one is identified from surrounding context. A scan may return an institution that appears in the filing for reasons unrelated to the offering, in which case the search is bounded to the sections where underwriters are named.

Each of these is demonstrated later in this methodology on real filings, including one case where the naive answer is wrong and the reason is visible in the output.

Why Standard IPO Analysis Misses the Structure

Even an analyst who ignores the first-day move and goes to the offering terms can be misled. Four failure modes recur.

The 7% Anchoring Trap

The underwriting spread looks like the most obvious structural measure, and read in isolation it is easy to misinterpret. A widely cited figure holds that US IPOs price at a 7% gross spread. That figure describes smaller offerings well and larger ones poorly. Spreads fall as offerings grow, because scale is associated with lower spreads, although size is one contributing factor among several including complexity, risk, expected demand, and underwriter competition.

The consequence is that a spread carries information only when it is read against the size of the offering. A mid-size deal at the rate a small deal would pay is worth noticing. A large deal at a mid-size rate is worth noticing. The absolute number, compared across offerings of different scale, is not.

The Structured-Data Ceiling

The second trap is treating the fields a data source exposes as the whole picture. The cover-page economics are easy to work with, and an analysis can be built on them quickly. But the syndicate, the share counts, and the primary and secondary split are absent from those fields, and an analysis confined to them cannot see any of the three. They are recoverable only from the filing.

Cash Spread and Selling-Shareholder Blind Spots

Two distortions hide inside offerings that look ordinary. The cash spread can understate what the underwriters received, because some offerings grant additional non-cash compensation such as warrants, which this methodology records separately rather than folding into the spread. And an offering can be substantially a sale by existing shareholders rather than a capital raise by the company, which the headline offering size does not distinguish.

A related trap sits inside the second point. Selling-shareholder language appears in most prospectuses regardless of whether the base offering includes any secondary shares, because existing holders frequently grant the over-allotment option without selling in the offering itself. Detecting the phrase is not the same as measuring the split.

The Template Assumption

The final trap is assuming every offering fits a standard template. Many depart from it. A company may have multiple share classes with different economic rights, so an offered-share ratio depends entirely on which denominator is chosen. An Up-C structure holds part of the economic interest in units exchangeable into listed shares rather than in the listed class itself. An over-allotment option may be granted by existing holders rather than by the company, so an identical share count can mean different things in two offerings. A methodology that assumes one structure produces confident numbers that answer the wrong question.

These four are independent. Handling one does not address the others, and the design that follows addresses each explicitly.

What Has to Be Engineered

Constructing the IPO Universe

The universe is defined by rule, so that it can be rebuilt identically. This methodology includes offerings that are marked priced in the calendar, list on a major US exchange, are issued by an operating company rather than a blank-check vehicle, fund, or trust, and carry a common-stock ticker rather than a unit, warrant, or rights tranche. The base offering excluding the over-allotment option is the consistent basis for offering size, the offered-share numerator, and the primary and secondary split.

Two of those screens are applied from structured fields, and two from documented heuristics on the company name and ticker form. Heuristics are stated as such, because a name-based test will occasionally misclassify an issuer, and a reader reproducing the universe should know which filters are exact and which are approximate.

Normalizing Structured Fields and Reconciling Coverage

Coverage of the structured economics is partial, and the endpoint that carries them returns prospectus filings generally rather than IPO prospectuses only, so a company that has raised capital since listing appears more than once. Both conditions are handled by matching records to the pricing date and by treating the absence of a record as a normal condition in which the economics are read from the filing instead.

There is one further normalization the data does not settle. The offering total and the calendar share count do not always reconcile against the offering price, which means the treatment of the over-allotment option is not consistent across records. Where the two disagree, the filing decides.

Retrieving the Prospectus Reliably

Locating the document is a retrieval problem with two specific requirements. Filings are requested by form type, and the response is capped, so the request has to be paged until a short page returns. And because a company's filing history includes later offerings, the filing is treated as the IPO prospectus only when it is dated near the pricing date.

Applied to a full year of offerings, this returns a final prospectus for most of the universe, and the overwhelming majority are 424B4 filings, with 424B1 and 424B3 appearing occasionally.

Validating What the Scan Returns

Extraction produces candidates. Validation turns them into measurements, and the methodology treats the two as different things throughout. A share-count scan returns several values, of which one is the post-offering count for the relevant class and the others are pre-offering counts, other classes, or figures that assume the over-allotment is exercised. An institution scan over a whole filing returns banks that appear for reasons unrelated to the offering.

Both are addressed in the process rather than in a caveat. Share-count candidates are returned with the surrounding text so the correct one can be identified. The institution search is bounded to the cover page and the underwriting section. Anything that cannot be resolved confidently is flagged for review rather than reported.

How FMP Supports the Workflow

Defining the Universe With the IPOs Calendar API

The IPOs Calendar API supplies the population: offerings with their pricing dates, exchange, status, company name, and share counts. It is the starting point for every screen in the universe rule.

One property of the endpoint shapes how it is queried. The number of rows returned falls as the requested date range widens, so a single request covering several months returns materially fewer offerings than the same period requested month by month. The universe is therefore assembled from month-sized windows, and the walkthrough shows the difference directly.

Calculating Offering Economics With the IPOs Prospectus API

The IPOs Prospectus API returns the cover-page economics as structured fields: public offering price per share and in total, underwriting discounts and commissions, and proceeds before expenses, along with filing dates, the CIK, and a link to the filing. The gross spread and the retained fraction are computed directly from these.

The endpoint carries prospectus filings generally rather than IPO prospectuses only, so records are matched to the pricing date before use, and offerings without a record are read from the filing instead.

Locating Prospectus Filings With the SEC Filings Search API

Locating each offering's final prospectus is done through the SEC filings search endpoint, queried by form type for 424B4, 424B1, and 424B3. It returns the filing date, the form type, and a direct link to the document itself, which is what the scan retrieves.

Two requirements apply. The response is capped per request, so queries are paged. And because the response includes later offerings by the same companies, the result is matched to the pricing date before it is treated as the IPO prospectus.

Dividing Responsibilities Between FMP and the Methodology

FMP supplies the population, the structured economics where they exist, and a reliable path to the filing. The methodology supplies the universe rule, the date matching, the document scan, the validation, and the structural read. The SEC filing remains the authoritative source for every reported figure, and the structured fields serve as a cross-check against it.

That division keeps the engineering effort in the analysis rather than in assembling and hosting IPO and filing data, which is the larger problem.

Framework Design

The methodology starts from the terms disclosed on the cover of the prospectus and derives measures from them. There is no composite score. Each measure is reported separately, alongside a note of what was available for that offering, because combining measures into a single figure would require weights this methodology does not attempt to justify.

The Base Unit

The base unit is the set of disclosed terms: the offering price per share, the underwriting discount per share, the proceeds before expenses, the total offering amount, and the share counts. These are facts rather than measures. Everything below is a transformation of them, and the base unit is the input to the methodology rather than one of its measures.

Gross Spread

The gross spread is the underwriting discounts and commissions per share divided by the public offering price per share. It covers cash consideration only. Where a prospectus discloses non-cash compensation such as underwriter warrants, that is recorded separately and never folded into this figure, because doing so would make the measure incomparable across offerings that compensate differently.

The spread is read against the size of the offering rather than in absolute terms, for the reason set out under Why Standard IPO Analysis Misses the Structure.

The Size Benchmark

A spread carries information only against the size of the offering, so each spread is compared to a benchmark for offerings of similar size. The benchmark is fixed in advance rather than derived from the quarter being studied, so that a spread means the same thing from one quarter to the next.

Benchmarks are median gross spreads by size band, computed over the eight quarters ending at the last benchmark period, from this methodology's own universe restricted to offerings with a gross spread drawn from a matched structured record. The cover-page pricing table and the structured record report the same figures, so a matched record is a sufficient basis for a population benchmark, while any spread reported for an individual offering is validated against the filing before publication. The six bands split base offering size at fifty million, one hundred million, two hundred fifty million, five hundred million, and one billion dollars. A band is benchmarked only when it holds at least twenty such offerings; thinner bands are reported as unbenchmarked rather than merged into their neighbours, and offerings in them carry an absolute spread and no deviation.

Studies report the absolute spread and its deviation in percentage points from the applicable benchmark. The deviation is a distance, not a verdict, and the benchmark controls for nothing beyond size. The benchmark table is published separately and versioned, and each study names the version it applied.

Retained Fraction

The retained fraction is the proceeds before expenses per share divided by the offering price per share. It is the arithmetic complement of the gross spread and carries no independent information, so it is displayed beside the spread rather than reported as a separate measure. It is included because the share of each share's price that remains after the discount is often the more intuitive way to read the same fact.

It is a per-share measure. It is not the fraction of the total raise the company keeps, which also depends on whether existing shareholders sold in the offering.

Underwriter Syndicate

The syndicate is recorded as observable fact: which institutions are named as underwriters on the cover of the prospectus, and in what roles. This methodology does not assign banks a quality tier, because a ranking is an editorial judgment rather than a disclosed fact.

The automated scan produces a candidate list, not a syndicate. Institutions are confirmed from the cover page and the underwriting section before being reported, and the count reported is the count the filing states.

Offered-Share Ratio

The offered-share ratio is the shares offered in the base offering measured against the shares outstanding after the offering. It requires a denominator decision, and this methodology reports the decision explicitly rather than burying it.

  • Single class: shares outstanding after the offering.
  • Multiple classes: the ratio is reported against the listed class and against all shares outstanding, because the two answer different questions and differ substantially.
  • Up-C structures: against the listed class, and against the fully exchanged base including units exchangeable into listed shares, with non-economic voting shares excluded.

This is an offered-share ratio. It is not the SEC's regulatory public float, which is based on the market value of equity held by non-affiliates and is computed differently. And it is not a measure of future tradable supply, because lockups, registration rights, conversion terms, and holder intent all sit between a share count and a share that trades.

Selling-Shareholder Ratio

The selling-shareholder ratio is the shares offered by existing holders divided by the total shares in the base offering. Those sellers may be founders, employees, funds, or other existing holders, so the measure describes the split between company shares and existing-shareholder shares rather than characterizing who is selling.

The measure is read from the offering-share table, not inferred from the presence of selling-shareholder language, for the reason given under Why Standard IPO Analysis Misses the Structure. No direction is assigned. A larger secondary component is neither favorable nor unfavorable on its own.

What the Methodology Produces

Every offering in every quarterly refresh is recorded in the same schema, so that studies are comparable and a reader can see what was measured and how it was obtained. The required fields are:

  • Company name and ticker
  • Pricing date
  • Base offering size, excluding the over-allotment option
  • Offering price per share
  • Gross spread, and its deviation from the applicable size benchmark, or a note that the band is unbenchmarked
  • Retained fraction
  • Syndicate: lead bookrunners, other named underwriters, and total syndicate count
  • Offered-share ratio, reported once for each applicable denominator, with each denominator named
  • Primary and secondary share counts
  • Selling-shareholder ratio
  • Non-cash underwriter compensation, where the prospectus discloses it
  • Source type for each measure: structured field, validated reading of the filing, or both
  • Validation status for each measure
  • Benchmark version applied

Validation status distinguishes three conditions that are easy to conflate. A measure is validated when its value was confirmed against the filing. It is unvalidated when a value was retrieved but could not be confirmed, which is a different thing from having no value at all. It is not disclosed when the offering's documents do not contain it. An unvalidated figure is never reported as a measurement and never enters an aggregate.

Syndicate categories are fixed across studies so that counts mean the same thing. Lead bookrunners are the institutions the prospectus cover names in that role. Other named underwriters are the remaining institutions on the cover. Total syndicate count is the number of institutions the cover names, which is taken from the filing rather than from the scan.

There is no combined score, and comparisons between offerings are made measure by measure.

Reproducing the Workflow With FMP

This walkthrough runs the process end to end. Each stage states the decision it implements, gives the code, shows the output, and interprets the result. The universe is built for a full calendar year, and two offerings are then read in detail: a very large offering and a conventional mid-size one, chosen because their share structures differ in ways that exercise the parts of the methodology most easily got wrong, above all the denominator question that multi-class and Up-C structures force.

Before You Run This Code

You will need Python 3.9 or later, the requests library, an FMP API key, and a contact email, since the SEC requires a descriptive User-Agent header on document requests. The full script appears under Complete Reference Implementation.

import requests, re, time

from datetime import datetime


API_KEY = "YOUR_FMP_API_KEY"

SEC_EMAIL = "your-email@example.com"


BASE = "https://financialmodelingprep.com/stable"

SEC_HEADERS = {"User-Agent": "Signals Lab Research " + SEC_EMAIL}


US_EXCHANGES = {"NASDAQ", "NYSE", "AMEX", "NYSE AMERICAN"}

NON_OPERATING = ("acquisition corp", "acquisition co", "acquisition holdings",

" spac", "blank check", "trust", "fund", "etf", " reit")


def api(endpoint, params):

p = dict(params); p["apikey"] = API_KEY

r = requests.get(BASE + "/" + endpoint, params=p, timeout=60)

r.raise_for_status()

d = r.json()

if not isinstance(d, list):

raise RuntimeError("unexpected response from %s" % endpoint)

return d


def to_number(x):

if x is None or x == "":

return None

try:

return float(str(x).replace(",", ""))

except ValueError:

return None


def to_date(x):

try:

return datetime.strptime(str(x)[:10], "%Y-%m-%d")

except Exception:

return None


def is_operating_company(name):

n = str(name or "").lower()

return not any(w in n for w in NON_OPERATING)


def is_common_stock_ticker(sym):

return bool(sym) and len(sym) <= 4 and sym.isalpha()

Stage 1: Build the Universe

The universe comes from the IPOs Calendar API, requested month by month for the reason described under How FMP Supports the Workflow, then filtered by the universe rule.

def monthly_windows(year):

"""The calendar returns fewer rows as the requested range widens, so the

universe is assembled from month-sized windows."""

last = [31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31]

if year % 4 == 0 and (year % 100 != 0 or year % 400 == 0):

last[1] = 29

return [("%d-%02d-01" % (year, m), "%d-%02d-%02d" % (year, m, last[m-1]))

for m in range(1, 13)]


def build_universe(windows):

universe = {}

for start, end in windows:

for row in api("ipos-calendar", {"from": start, "to": end}):

sym = row.get("symbol")

if not sym or "." in sym:

continue

if str(row.get("exchange") or "").upper() not in US_EXCHANGES:

continue

if str(row.get("actions") or "").strip().lower() != "priced":

continue

if not is_operating_company(row.get("company")):

continue

if not is_common_stock_ticker(sym):

continue

priced = to_date(row.get("date"))

if priced is None:

continue

prior = universe.get(sym)

if prior is None or (priced and prior["priced"] and priced < prior["priced"]):

universe[sym] = {"priced": priced,

"company": row.get("company"),

"shares_offered": to_number(row.get("shares"))}

return universe

Output: Stage 1, the window comparison and the resulting universe.

The first two lines are the reason for the monthly loop. The same six-month period returns substantially fewer rows when requested in a single call than when requested month by month, and nothing in the response indicates that anything is missing. A pipeline that requests wide ranges silently analyzes a fraction of the market. After the universe rule is applied, the result is the set of priced US common-stock IPOs by operating companies for the year.

Stage 2: Locate Each Prospectus

Each offering's final prospectus is located through the SEC filings search endpoint, paged because the response is capped, and matched to the pricing date because a company's filing history includes its later offerings.

PROSPECTUS_FORMS = ("424B4", "424B1", "424B3")


def fetch_prospectus_filings(windows, forms=PROSPECTUS_FORMS):

"""The filings endpoint returns at most 100 rows per call, so each form

type and window is paged until a short page comes back."""

filings = {}

for form in forms:

for start, end in windows:

page = 0

while page < 10:

rows = api("sec-filings-search/form-type",

{"formType": form, "from": start, "to": end,

"page": page, "limit": 1000})

if not rows:

break

for row in rows:

sym = row.get("symbol")

if sym and sym != "None":

filings.setdefault(sym, []).append(row)

if len(rows) < 100:

break

page += 1

time.sleep(0.2)

return filings


def match_prospectus(universe, filings, max_gap_days=45):

"""A company's prospectus filings include its later offerings, so a filing

counts as the IPO prospectus only when it sits near the pricing date."""

for sym, deal in universe.items():

best = None

for row in filings.get(sym, []):

filed = to_date(row.get("filingDate"))

if not filed or not deal["priced"]:

continue

gap = abs((filed - deal["priced"]).days)

if best is None or gap < best["gap"]:

best = {"gap": gap, "form": row.get("formType"),

"document": row.get("finalLink")}

deal["prospectus"] = best if (best and best["gap"] <= max_gap_days) else None

return universe

Output: Stage 2, matched prospectuses and their form types.

Most of the universe resolves to a final prospectus, and the form distribution confirms that the 424B4 is the filing that matters, with 424B1 and 424B3 appearing in a handful of cases. The date match is what makes this reliable. Without it, a company that has raised capital since listing contributes its most recent offering rather than its IPO.

Stage 3: Attach Structured Economics

Where the IPOs Prospectus API has a record for an offering, the gross spread and retained fraction follow from the cover-page fields. Records are matched to the pricing date for the same reason the filings are.

def fetch_structured(windows):

records = {}

for start, end in windows:

for row in api("ipos-prospectus", {"from": start, "to": end}):

sym = row.get("symbol")

if sym:

records.setdefault(sym, []).append(row)

return records


def offering_economics(deal, records):

best = None

for row in records:

price = to_number(row.get("pricePublicPerShare"))

discount = to_number(row.get("discountsAndCommissionsPerShare"))

proceeds = to_number(row.get("proceedsBeforeExpensesPerShare"))

total = to_number(row.get("pricePublicTotal"))

if not price or price <= 0 or discount is None or total is None:

continue

filed = to_date(row.get("filingDate"))

gap = abs((filed - deal["priced"]).days) if (filed and deal["priced"]) else 9999

if best is None or gap < best["gap"]:

best = {"gap": gap, "price": price, "total": total,

"gross_spread": discount / price,

"retained": (proceeds / price) if proceeds is not None else None}

return best if (best and best["gap"] <= 45) else None

Output: Stage 3, offerings with a structured economics record.

Structured economics are available for a minority of offerings rather than for the whole universe, which the workflow treats as an ordinary condition. An offering without a record is not dropped. Its economics are read from the cover page of the filing instead, which is the authoritative source in either case.

Stage 4: Scan the Prospectus

The document is retrieved and scanned for candidate facts. Two design decisions are visible in the code: the institution search is bounded to the sections where underwriters are named, and share-count matches are returned with their surrounding text.

BANK_ALIASES = {

"Goldman Sachs": ["Goldman Sachs"], "Morgan Stanley": ["Morgan Stanley"],

"BofA": ["BofA", "Merrill"], "Citigroup": ["Citigroup"],

"J.P. Morgan": ["J.P. Morgan", "JPMorgan"], "Barclays": ["Barclays"],

"Deutsche Bank": ["Deutsche Bank"], "RBC": ["RBC"], "UBS": ["UBS"],

"Wells Fargo": ["Wells Fargo"], "Jefferies": ["Jefferies"],

"Cantor": ["Cantor"], "Evercore": ["Evercore"], "Needham": ["Needham"],

"Raymond James": ["Raymond James"], "Stifel": ["Stifel"],

"William Blair": ["William Blair"], "Mizuho": ["Mizuho"],

"Santander": ["Santander"], "BTG Pactual": ["BTG Pactual"],

"Macquarie": ["Macquarie"], "Societe Generale": ["Societe Generale"],

}


AFTER_PHRASES = [r"outstanding immediately after th(?:is|e) offering",

r"outstanding after th(?:is|e) offering",

r"to be outstanding after th(?:is|e) offering"]


def plain_text(url):

r = requests.get(url, headers=SEC_HEADERS, timeout=90)

r.raise_for_status()

t = re.sub(r"<[^>]+>", " ", r.text)

t = re.sub(r" | ", " ", t)

return re.sub(r"\s+", " ", t)


def underwriting_region(text):

"""Cover page plus the underwriting section. Institution names elsewhere

in the filing are usually lenders, advisers or counterparties."""

region = text[:60000]

for m in re.finditer(r"\bUNDERWRITING\b", text):

region += " " + text[m.start(): m.start() + 60000]

return region


def institutions_named(text):

return [f for f, spellings in BANK_ALIASES.items()

if any(s in text for s in spellings)]


def share_count_candidates(text, window=140):

"""Share counts near a post-offering phrase, returned with surrounding

words so the reader can tell which class and which basis each refers to."""

found, seen = [], set()

for phrase in AFTER_PHRASES:

for m in re.finditer(phrase, text, re.IGNORECASE):

lo, hi = max(0, m.start() - window), min(len(text), m.end() + window)

context = text[lo:hi]

for num in re.finditer(r"\b([\d]{1,3}(?:,[\d]{3}){2,})\b", context):

value = int(num.group(1).replace(",", ""))

if value in seen:

continue

seen.add(value)

found.append({"formatted": num.group(1), "value": value,

"context": context.strip()})

return sorted(found, key=lambda x: -x["value"])


def scan_prospectus(url):

text = plain_text(url)

return {"characters": len(text),

"institutions_document": institutions_named(text),

"institutions_underwriting": institutions_named(underwriting_region(text)),

"share_candidates": share_count_candidates(text),

"selling_shareholder_language": ("selling stockholder" in text.lower()

or "selling shareholder" in text.lower())}

Output: Stage 4, candidate facts from both prospectuses.

The two filings illustrate opposite outcomes, which is why both are shown. For the very large offering, bounding the search changes nothing: every institution named anywhere in the filing is also named in the underwriting section. For the mid-size offering, bounding removes two institutions that appear elsewhere in the document but are not underwriters of the deal. A pipeline that scanned the whole filing would report a syndicate two firms larger than the one on the cover.

The comparison also shows what a name-list scan cannot do. For the large offering the scan reports twenty institutions, while the cover names twenty-two, because two underwriters are absent from the reference list. A scan can return names that are not underwriters and can miss underwriters that are, which is why the reported count comes from the filing rather than from the scan.

The share-count candidates make the validation problem concrete. For the large offering, four values appear: the shares outstanding after the offering for the listed class, the same figure assuming the over-allotment is exercised, the count for the other class, and the pre-offering count for the listed class. The pre-offering figure is the trap. It sits adjacent to a post-offering phrase, it is the right order of magnitude, and a scan that takes the first number following the phrase will select it and produce a ratio that is not a ratio of anything. The context returned with each candidate is what makes the difference visible.

Stage 5: Report the Structural Read

The final stage reports the measures from figures confirmed against the filing. This is a deliberate separation. The scan proposes; the filing decides.

def structural_read(name, offered, denominators, secondary_shares, syndicate):

"""Reported from figures confirmed against the filing, not from the scan."""

if not offered:

raise ValueError("shares offered must be positive")

print(" %s" % name)

print(" shares offered : %s" % format(offered, ","))

for label, denom in denominators:

print(" offered / %-22s: %6.2f%% (denominator %s)"

% (label, offered / denom * 100, format(denom, ",")))

print(" secondary shares offered : %s (%.2f%% of the offering)"

% (format(secondary_shares, ","), secondary_shares / offered * 100))

print(" underwriters on the cover : %d" % syndicate)

Output: Stage 5, the structural read for both offerings.

Both offerings report two offered-share ratios, and the gap between them is the point. For the large offering the ratio against the listed class is roughly seven and a half percent, while the ratio against all shares outstanding is a little over four percent, because a second class holds a large share of the economic interest. For the mid-size offering, an Up-C structure, the ratio against the listed class is above thirty-seven percent, while the ratio against the fully exchanged base is around twenty-three percent, because a substantial portion of the economic interest sits in exchangeable units rather than in listed shares.

Neither number is more correct than the other. They answer different questions, and reporting only one without naming the denominator would be misleading in both cases. Both offerings also report a secondary ratio of zero: in each, the company offered every share in the base offering, and where an existing holder is involved it granted the over-allotment option rather than selling in the offering itself. That is the distinction drawn earlier under Why Standard IPO Analysis Misses the Structure, and it is why the ratio is read from the offering-share table rather than inferred from the presence of selling-shareholder language.

Complete Reference Implementation

The script below runs the full workflow. It builds the universe for a calendar year, locates each prospectus, attaches structured economics where they exist, scans two documents, and reports the structural read from validated figures. It is written for clarity rather than production deployment.

Two behaviors are deliberate and should be preserved in any adaptation. The scan returns candidates rather than conclusions, and the structural read is driven by figures confirmed against the filing rather than by the scan output.

"""

IPO Offering Structure: reference implementation.


Stage 1 Build the universe from the IPO calendar.

Stage 2 Locate each deal's final prospectus.

Stage 3 Attach structured offering economics where a record exists.

Stage 4 Scan the prospectus for candidate facts.

Stage 5 Report the structural read from validated figures.

"""


import requests, re, time

from datetime import datetime


API_KEY = "YOUR_FMP_API_KEY"

SEC_EMAIL = "your-email@example.com"


BASE = "https://financialmodelingprep.com/stable"

SEC_HEADERS = {"User-Agent": "Signals Lab Research " + SEC_EMAIL}


US_EXCHANGES = {"NASDAQ", "NYSE", "AMEX", "NYSE AMERICAN"}

NON_OPERATING = ("acquisition corp", "acquisition co", "acquisition holdings",

" spac", "blank check", "trust", "fund", "etf", " reit")

PROSPECTUS_FORMS = ("424B4", "424B1", "424B3")



# ---------------------------------------------------------------- helpers

def api(endpoint, params):

p = dict(params); p["apikey"] = API_KEY

r = requests.get(BASE + "/" + endpoint, params=p, timeout=60)

r.raise_for_status()

d = r.json()

if not isinstance(d, list):

raise RuntimeError("unexpected response from %s" % endpoint)

return d



def to_number(x):

if x is None or x == "":

return None

try:

return float(str(x).replace(",", ""))

except ValueError:

return None



def to_date(x):

try:

return datetime.strptime(str(x)[:10], "%Y-%m-%d")

except Exception:

return None



def monthly_windows(year):

"""The calendar returns fewer rows as the requested range widens, so the

universe is assembled from month-sized windows."""

last = [31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31]

if year % 4 == 0 and (year % 100 != 0 or year % 400 == 0):

last[1] = 29

return [("%d-%02d-01" % (year, m), "%d-%02d-%02d" % (year, m, last[m-1]))

for m in range(1, 13)]



# ------------------------------------------------- Stage 1: the universe

def is_operating_company(name):

n = str(name or "").lower()

return not any(w in n for w in NON_OPERATING)



def is_common_stock_ticker(sym):

return bool(sym) and len(sym) <= 4 and sym.isalpha()



def build_universe(windows):

universe = {}

for start, end in windows:

for row in api("ipos-calendar", {"from": start, "to": end}):

sym = row.get("symbol")

if not sym or "." in sym:

continue

if str(row.get("exchange") or "").upper() not in US_EXCHANGES:

continue

if str(row.get("actions") or "").strip().lower() != "priced":

continue

if not is_operating_company(row.get("company")):

continue

if not is_common_stock_ticker(sym):

continue

priced = to_date(row.get("date"))

if priced is None:

continue

prior = universe.get(sym)

if prior is None or (priced and prior["priced"] and priced < prior["priced"]):

universe[sym] = {"priced": priced,

"company": row.get("company"),

"shares_offered": to_number(row.get("shares"))}

return universe



# ------------------------------------------ Stage 2: locate the prospectus

def fetch_prospectus_filings(windows, forms=PROSPECTUS_FORMS):

"""The filings endpoint returns at most 100 rows per call, so each form

type and window is paged until a short page comes back."""

filings = {}

for form in forms:

for start, end in windows:

page = 0

while page < 10:

rows = api("sec-filings-search/form-type",

{"formType": form, "from": start, "to": end,

"page": page, "limit": 1000})

if not rows:

break

for row in rows:

sym = row.get("symbol")

if sym and sym != "None":

filings.setdefault(sym, []).append(row)

if len(rows) < 100:

break

page += 1

time.sleep(0.2)

return filings



def match_prospectus(universe, filings, max_gap_days=45):

"""A company's prospectus filings include its later offerings, so a filing

counts as the IPO prospectus only when it sits near the pricing date."""

for sym, deal in universe.items():

best = None

for row in filings.get(sym, []):

filed = to_date(row.get("filingDate"))

if not filed or not deal["priced"]:

continue

gap = abs((filed - deal["priced"]).days)

if best is None or gap < best["gap"]:

best = {"gap": gap, "form": row.get("formType"),

"document": row.get("finalLink")}

deal["prospectus"] = best if (best and best["gap"] <= max_gap_days) else None

return universe



# ------------------------------------------- Stage 3: structured economics

def fetch_structured(windows):

records = {}

for start, end in windows:

for row in api("ipos-prospectus", {"from": start, "to": end}):

sym = row.get("symbol")

if sym:

records.setdefault(sym, []).append(row)

return records



def offering_economics(deal, records):

best = None

for row in records:

price = to_number(row.get("pricePublicPerShare"))

discount = to_number(row.get("discountsAndCommissionsPerShare"))

proceeds = to_number(row.get("proceedsBeforeExpensesPerShare"))

total = to_number(row.get("pricePublicTotal"))

if not price or price <= 0 or discount is None or total is None:

continue

filed = to_date(row.get("filingDate"))

gap = abs((filed - deal["priced"]).days) if (filed and deal["priced"]) else 9999

if best is None or gap < best["gap"]:

best = {"gap": gap, "price": price, "total": total,

"gross_spread": discount / price,

"retained": (proceeds / price) if proceeds is not None else None}

return best if (best and best["gap"] <= 45) else None



# -------------------------------------------- Stage 4: scan the prospectus

BANK_ALIASES = {

"Goldman Sachs": ["Goldman Sachs"], "Morgan Stanley": ["Morgan Stanley"],

"BofA": ["BofA", "Merrill"], "Citigroup": ["Citigroup"],

"J.P. Morgan": ["J.P. Morgan", "JPMorgan"], "Barclays": ["Barclays"],

"Deutsche Bank": ["Deutsche Bank"], "RBC": ["RBC"], "UBS": ["UBS"],

"Wells Fargo": ["Wells Fargo"], "Jefferies": ["Jefferies"],

"Cantor": ["Cantor"], "Evercore": ["Evercore"], "Needham": ["Needham"],

"Raymond James": ["Raymond James"], "Stifel": ["Stifel"],

"William Blair": ["William Blair"], "Mizuho": ["Mizuho"],

"Santander": ["Santander"], "BTG Pactual": ["BTG Pactual"],

"Macquarie": ["Macquarie"], "Societe Generale": ["Societe Generale"],

}


AFTER_PHRASES = [r"outstanding immediately after th(?:is|e) offering",

r"outstanding after th(?:is|e) offering",

r"to be outstanding after th(?:is|e) offering"]



def plain_text(url):

r = requests.get(url, headers=SEC_HEADERS, timeout=90)

r.raise_for_status()

t = re.sub(r"<[^>]+>", " ", r.text)

t = re.sub(r" | ", " ", t)

return re.sub(r"\s+", " ", t)



def underwriting_region(text):

"""Cover page plus the underwriting section. Institution names elsewhere

in the filing are usually lenders, advisers or counterparties."""

region = text[:60000]

for m in re.finditer(r"\bUNDERWRITING\b", text):

region += " " + text[m.start(): m.start() + 60000]

return region



def institutions_named(text):

return [f for f, spellings in BANK_ALIASES.items()

if any(s in text for s in spellings)]



def share_count_candidates(text, window=140):

"""Share counts near a post-offering phrase, returned with surrounding

words so the reader can tell which class and which basis each refers to."""

found, seen = [], set()

for phrase in AFTER_PHRASES:

for m in re.finditer(phrase, text, re.IGNORECASE):

lo, hi = max(0, m.start() - window), min(len(text), m.end() + window)

context = text[lo:hi]

for num in re.finditer(r"\b([\d]{1,3}(?:,[\d]{3}){2,})\b", context):

value = int(num.group(1).replace(",", ""))

if value in seen:

continue

seen.add(value)

found.append({"formatted": num.group(1), "value": value,

"context": context.strip()})

return sorted(found, key=lambda x: -x["value"])



def scan_prospectus(url):

text = plain_text(url)

return {"characters": len(text),

"institutions_document": institutions_named(text),

"institutions_underwriting": institutions_named(underwriting_region(text)),

"share_candidates": share_count_candidates(text),

"selling_shareholder_language": ("selling stockholder" in text.lower()

or "selling shareholder" in text.lower())}



# ------------------------------------------------------------- Stage 5

def structural_read(name, offered, denominators, secondary_shares, syndicate):

"""Reported from figures confirmed against the filing, not from the scan."""

if not offered:

raise ValueError("shares offered must be positive")

print(" %s" % name)

print(" shares offered : %s" % format(offered, ","))

for label, denom in denominators:

print(" offered / %-22s: %6.2f%% (denominator %s)"

% (label, offered / denom * 100, format(denom, ",")))

print(" secondary shares offered : %s (%.2f%% of the offering)"

% (format(secondary_shares, ","), secondary_shares / offered * 100))

print(" underwriters on the cover : %d" % syndicate)



# ------------------------------------------------------------- run

def run():

year = 2025

windows = monthly_windows(year)


print("=" * 68)

print("STAGE 1 Build the universe")

print("=" * 68)

one_call = api("ipos-calendar", {"from": "%d-01-01" % year, "to": "%d-06-30" % year})

per_month = []

for s, e in windows[:6]:

per_month += api("ipos-calendar", {"from": s, "to": e})

print(" first half of %d requested in one call : %d rows" % (year, len(one_call)))

print(" first half of %d requested month by month: %d rows" % (year, len(per_month)))

universe = build_universe(windows)

print(" priced US common-stock IPOs of operating companies in %d: %d"

% (year, len(universe)))


print("\n" + "=" * 68)

print("STAGE 2 Locate each prospectus")

print("=" * 68)

filings = fetch_prospectus_filings(windows)

universe = match_prospectus(universe, filings)

matched = [d for d in universe.values() if d.get("prospectus")]

forms = {}

for d in matched:

forms[d["prospectus"]["form"]] = forms.get(d["prospectus"]["form"], 0) + 1

print(" deals with a prospectus filed near pricing: %d" % len(matched))

print(" form types of those filings : %s" % forms)


print("\n" + "=" * 68)

print("STAGE 3 Attach structured offering economics")

print("=" * 68)

records = fetch_structured(windows)

have = 0

for sym, deal in universe.items():

deal["economics"] = offering_economics(deal, records.get(sym, []))

if deal["economics"]:

have += 1

print(" deals with a structured economics record: %d" % have)

print(" the endpoint also carries later offerings, so records are matched")

print(" on the pricing date the same way the filings are.")


print("\n" + "=" * 68)

print("STAGE 4 Scan the prospectus for candidate facts")

print("=" * 68)

demos = {"SPCX": ("https://www.sec.gov/Archives/edgar/data/1181412/"

"000162828026042639/spaceexplorationtechnologi.htm"),

"SUJA": ("https://www.sec.gov/Archives/edgar/data/1934114/"

"000110465926057504/tm2530822-13_424b4.htm")}

for sym, url in demos.items():

scan = scan_prospectus(url)

doc, und = scan["institutions_document"], scan["institutions_underwriting"]

print("\n %s (%s characters)" % (sym, format(scan["characters"], ",")))

print(" institutions, whole document : %d" % len(doc))

print(" institutions, underwriting only : %d %s" % (len(und), und))

excluded = [b for b in doc if b not in und]

print(" removed by bounding the search : %s" % (excluded or "none"))

print(" share-count candidates:")

for c in scan["share_candidates"][:4]:

print(" %16s ...%s..." % (c["formatted"], c["context"][:88]))


print("\n" + "=" * 68)

print("STAGE 5 Structural read from validated figures")

print("=" * 68)

structural_read("SPCX Space Exploration Technologies",

offered=555555555,

denominators=[("Class A outstanding", 7380196910),

("all shares outstanding", 13075865175)],

secondary_shares=0, syndicate=22)

print()

structural_read("SUJA Suja Life (Up-C structure)",

offered=8888889,

denominators=[("Class A outstanding", 23788700),

("fully exchanged base", 38625012)],

secondary_shares=0, syndicate=5)



if __name__ == "__main__":

run()



A production version would add caching of both API responses and retrieved documents, rate-limit handling, logging, a review queue for candidates that cannot be resolved automatically, and persistence so that a past period can be reproduced exactly. It would also record, for every reported figure, whether it came from a structured field or from a validated reading of the filing.

Running the Workflow at Scale

The Structured Stages

Building the universe and attaching economics are inexpensive. The cost is a fixed number of requests per month of coverage, and the results are stable once a period has closed, so each period is pulled once and cached. The work is reconciliation rather than volume: matching calendar entries to filings and to economics records, and recording which offerings have which.

The Document Stage

Retrieving and scanning documents is the expensive stage. Each filing is a large download and each scan produces candidates that require review. Documents are cached so that a filing is retrieved once, requests to the SEC respect its rate and identification requirements, and the review step is treated as part of the pipeline rather than as an afterthought.

Because validation is where analyst time goes, a deployment benefits from ordering the queue deliberately: the largest offerings first, then a rotating sample of the rest, so that the process does not systematically overlook the offerings nobody happened to select.

From Workflow to Dataset

Run over successive periods, the result is a dataset indexed by offering, carrying the structural measures and a record of how each was obtained. Because the universe rule and the measures are fixed, offerings remain comparable across periods, which is what makes questions about direction answerable at all.

Two constraints apply to those questions. Measures drawn from structured fields cover the offerings that have records, and measures drawn from documents cover the offerings that have been retrieved and validated. Any statement about a period should be scoped to the set it was computed from, and the methodology carries that scope alongside the numbers rather than dropping it.

Where the Methodology Breaks Down

Extraction Is Imperfect and Must Be Validated

The document stage is the most informative part of the methodology and the least reliable. Filings vary in structure and wording, and a scan tuned to one layout will miss or misread another. The walkthrough shows three failure modes on real filings: a share count adjacent to the right phrase that is nevertheless the wrong figure, institutions named in a filing that are not underwriters of the offering, and underwriters on the cover that a name-based scan does not recognize.

Neither error announces itself. A wrong number of the right magnitude looks exactly like a right one, which is why the methodology separates candidates from measurements and treats anything unresolved as unresolved rather than as a result.

Coverage Is Uneven

Structured economics exist for a minority of offerings, and documents are retrievable for most but not all of it. An offering can therefore carry a full set of measures, a partial set, or measures read entirely from its filing. The methodology records which, and any aggregate should state the set it was computed over. Treating a partial set as complete is the most likely way to produce a confident number that describes something other than the market.

Because coverage varies, every study reports it. Each discloses the number of eligible offerings in the quarter, how many were matched to a final prospectus, how many carried a structured economics record, how many were fully validated, and, for each reported measure, how many offerings produced a valid observation. These counts appear alongside the measures rather than in a footnote, because a median computed over nine offerings and one computed over ninety are different claims.

Two rules follow from this. A quarter-level figure is reported as a comparison only when valid observations cover at least half the eligible offerings for that measure; below that, the measure is labelled incomplete and excluded from aggregate interpretation, though the underlying values are still published. And aggregates are never assembled from figures that do not mean the same thing: prospectus-derived values are never imputed for offerings whose documents were not validated, and offered-share ratios computed on different denominators are never combined into a single number.

The Denominator Is a Choice

The offered-share ratio depends on which shares are counted, and for multi-class and Up-C structures the plausible denominators produce materially different answers, as the walkthrough shows. This methodology reports each ratio with its denominator named rather than selecting one and presenting it as the answer. That makes the numbers honest but it also means two offerings with different structures are not directly comparable on this measure without stating the basis.

Structural Facts Are Not Judgments

The methodology records what the filings disclose and stops there. A low spread is not evidence of a well-run offering; it may reflect scale. A broad syndicate is a fact about who underwrote the deal, not a verdict on it. A small offered-share ratio is a structural characteristic, not a warning. The step from a comparable fact to an investment view is analysis this methodology does not perform, and treating a structural measure as a rating is the most common way to misuse it.

These limitations come from a combination of the source disclosures, choices this methodology makes, and the demonstration implementation. Filing variability and uneven coverage are properties of the data. The denominator convention and the exclusion rules are methodology choices. The name-list scan and the phrase-based share-count search are limitations of the reference code, which a production version would improve.

Case Studies Using This Methodology

This page is the fixed reference for a quarterly series. Each quarter the framework is run across the full eligible universe for that quarter, not across a selection of offerings, and the resulting record is what the studies draw on. The universe rule, the measures, the retrieval process, and the validation requirements stay constant. What changes is the quarter.

Individual offerings are examined in more depth when the standardized scan flags them, against criteria fixed in advance: a gross spread far from the benchmark for its size band, an offered-share ratio at either extreme on a clearly stated denominator, a material selling-shareholder component, an unusual syndicate composition or disclosed non-cash compensation, or a disclosure structure that exposes a limitation of the methodology itself. These criteria identify offerings worth investigating. They are not a ranking, they do not imply that a flagged offering is better or worse than an unflagged one, and they say nothing about how any security will perform.

This methodology carries a version identifier and an effective date, both stated at the top of this page. Every study names the methodology version it applied, the quarter it covers, its data snapshot date, and the benchmark table version in force. A study without those four is not reproducible, because the same quarter run under different rules produces different numbers.

Changes are documented here before they are used. If a universe rule, benchmark definition, denominator convention, measure, validation threshold, or selection criterion changes, it is recorded on this page as a new version with its effective date, and it applies to studies published after that date. Recalculating the benchmark table on its schedule is a version change of the table, and studies name the table version they used.

Studies already published stay tied to the version they applied and are not silently recomputed under revised rules. A case study does not introduce a new measure or alter a convention to reach a conclusion; if something needs to change, it changes here first.

The first case study reads the largest offering in the walkthrough in full: what its spread indicates when read against its scale, what its syndicate and share structure disclose, and where the methodology's own limits appear, including the denominator question that a multi-class structure forces. Subsequent case studies apply the same process to later offerings.

For the method and the reference implementation, this page is the source. To see it applied, start with the case studies listed below.

Q2 2026: What the Offering Size Leaves Out: Reading Q2 2026 US IPOs From the Prospectus

About the Author
David Kirakosyan

Weekly Signals Desk analysis and API-driven market workflows

David Kirakosyan writes the Weekly Signals Desk for FMP, breaking down market signals while showing readers how to build similar workflows using the FMP API. His work focuses on turning raw API data into practical market analysis and repeatable workflows that developers and analysts can adapt to their own research.

Related

Financial data for every need

Real-time quotes and 30+ years of historical data, including prices, fundamentals, and insider transactions — all accessible via API.

Create Free Account