FMPFMP
Datasets
Insights/Data in Action/Screening Models/How to Construct Transparent Value, Momentum, and Quality Factors From Financial and Market Data

How to Construct Transparent Value, Momentum, and Quality Factors From Financial and Market Data

·

·20 min read
Data in Action

A factor model should make a large equity universe easier to compare without hiding the assumptions that drive the result. The challenge is not simply choosing familiar metrics. It is aligning data to one measurement date, defining every formula, controlling outliers and missing values, and making the final ranking reproducible.

This article builds an illustrative, researcher-created value, momentum, and quality model from FMP financial and market data.

Key Takeaways

  • A usable factor definition specifies its formula, direction, data window, units, treatment of invalid values, and rebalance timing before any stock is scored.
  • Winsorized percentile ranks place different inputs on a comparable scale. Tied observations receive average-rank scores, so the observed minimum and maximum do not necessarily equal 0 and 100.
  • Sector-neutral value and quality scores can reduce structural industry bias, while momentum may be more defensible across a broad liquid universe.
  • Historical tests require point-in-time membership and filing availability. Applying today's universe or today's TTM data to earlier dates introduces look-ahead and survivorship bias.

Why Transparency Matters in Factor Construction

Terms such as value, momentum, and quality sound precise, but each can represent several different calculations.

  • Value may mean earnings yield, free cash flow yield, or enterprise value relative to operating cash flow.
  • Momentum may use six months, twelve months, or a blended window that excludes the most recent month.
  • Quality may emphasize profitability, balance-sheet strength, cash conversion, or stability.

Two researchers can use the same labels and produce materially different rankings. The design choices behind thematic and factor-based investing determine what a score actually measures.

A transparent model resolves that ambiguity by publishing its universe, measurement date, formulas, normalization method, missing-data rules, weights, and rebalance schedule. The result becomes an auditable research output rather than a black box.

The framework below is intentionally illustrative. Its selected factors and weights are researcher-created choices, not FMP-endorsed settings or evidence of how FMP produces a rating or recommendation. The same discipline applies when scaling a fundamental or valuation screen across a larger coverage universe: the screening criteria, input definitions, and exclusions must remain visible.

Define the Universe and Measurement Date First

Assume a measurement cutoff of June 30, 2026, at 4:30 p.m. America/New_York time, equivalent to 20:30 UTC.

The historical universe is drawn from a dated security-universe snapshot containing securities active at that cutoff. It includes U.S.-listed common stocks with a historical market capitalization of at least $1 billion and a 20-day average daily dollar volume of at least $5 million. Funds, preferred shares, warrants, and securities inactive at the cutoff are excluded.

Companies that were active on June 30 but later acquired or delisted remain eligible when the historical dataset and research license support them.

Companies classified as Financial Services through the Company Profile API are excluded from this illustrative model. Banks, insurers, brokers, asset managers, consumer-finance companies, and other financial institutions have capital structures and operating measures that are not directly comparable with non-financial businesses under this specification. A model that includes financial companies should use separately defined inputs, such as price-to-book, return on equity, capital adequacy, and industry-specific solvency measures.

For the historical example:

  • The $1 billion threshold uses the marketCap field from the Historical Market Capitalization API for June 30, 2026.
  • The liquidity threshold is the arithmetic mean of close × volume across the 20 trading sessions ending on June 30, using the Historical Price EOD API.
  • The current Company Screener API and Company Profile API are useful for live-screening work, but their current active-security lists, market capitalizations, and classifications should not be used to reconstruct historical universe membership.

If a dated sector classification is unavailable and a current classification is used as a fallback, that limitation should be disclosed.

Measurement-date enterprise value is calculated as:

EV(t) = marketCap(t) + totalDebt(q) - cashAndCashEquivalents(q)

Here, marketCap(t) is the Historical Market Capitalization API value at the measurement date. totalDebt and cashAndCashEquivalents come from the latest eligible quarterly record in the Balance Sheet Statement API.

Apply a Conservative Availability Rule

An acceptedDate is compared with the 4:30 p.m. America/New_York cutoff only when the source provides an explicit UTC offset, identifies a time zone, or documents the time zone used for the field. The timestamp is then converted to America/New_York time before comparison.

When a source time zone is unknown, treat the timestamp as date-only information:

  • Records dated before June 30, 2026, are eligible.
  • Records dated after June 30, 2026, are ineligible.
  • Same-day records dated June 30, 2026, are ineligible because the model cannot establish that they were available before the cutoff.

The same conservative rule applies when only filingDate is available. Records dated June 29, 2026, or earlier are eligible; records dated June 30 or later are excluded.

An SEC acceptance timestamp is a useful availability control, but it does not independently establish the exact moment a document became publicly accessible or reached a data provider. EDGAR's public-dissemination process follows acceptance, so the model should not treat acceptance time as proof of immediate public or provider availability.

The cutoff also cannot prove that returned values are the original values available on that date. A provider may later replace an earlier record with restated or standardized figures. Point-in-time accuracy therefore requires preserved snapshots or values verified against the applicable filing or filing-specific record.

The audit trail should retain the query time, source endpoint, statement period, filing identity, availability fields, returned values, data-version information when available, and model version. If the original data version cannot be established, describe the result as a historical reconstruction and disclose possible restatement bias.

Choose Inputs With Clear Economic Roles

The model uses two inputs for each factor family. That is enough to demonstrate the construction process without allowing a large collection of correlated metrics to create false precision.

Family

Input

Illustrative formula

Better direction

Normalization group

Weight

Value

Free cash flow yield

TTM free cash flow ÷ market capitalization

Higher

Sector

20%

Value

Operating cash flow to enterprise value

TTM operating cash flow ÷ enterprise value

Higher

Sector

15%

Momentum

12-1 month return

adjClose(t-21) ÷ adjClose(t-252) - 1

Higher

Full universe

20%

Momentum

6-1 month return

adjClose(t-21) ÷ adjClose(t-126) - 1

Higher

Full universe

15%

Quality

Return on invested capital

TTM NOPAT ÷ average invested capital

Higher

Sector

20%

Quality

Net debt to EBITDA

Net debt ÷ TTM EBITDA

Lower

Sector

10%

The weights sum to 100%: value receives 35%, momentum 35%, and quality 30%. This allocation is an illustrative researcher-created choice for demonstrating the mechanics.

For the historical model, the financial inputs are constructed from eligible quarterly statements rather than copied from a current TTM endpoint. For a live screen, current values can be retrieved from the Key Metrics TTM API and Financial Ratios TTM API, timestamped at the time of the run. Those current values should not be presented as a reconstruction of the June 30, 2026 ranking.

Value Inputs

Free cash flow yield is defined as:

FCF Yield = TTM Free Cash Flow ÷ Measurement-Date Market Capitalization

Operating cash flow to enterprise value is defined as:

OCF/EV = TTM Operating Cash Flow ÷ Measurement-Date Enterprise Value

Both inputs are yields, so higher values receive better scores. A yield format avoids the discontinuity created by inverting a negative earnings multiple. Negative free cash flow or operating cash flow remains economically meaningful, so the model retains finite negative yields and controls their influence through winsorization.

For the June 30 historical ranking, TTM free cash flow and operating cash flow are constructed from eligible historical quarterly records from the Cash Flow Statement API.

Each TTM flow is the sum of four distinct, consecutive fiscal quarters using the freeCashFlow or operatingCashFlow field. Annual records must not be combined with quarterly records, and duplicate versions of the same fiscal quarter must not be counted twice. When multiple versions exist, use only a version that can be verified as available by the measurement cutoff. If four distinct consecutive eligible quarters cannot be established, mark the corresponding TTM input as missing.

OCF/EV is calculated only when enterprise value is positive. If enterprise value is zero, negative, missing, or non-finite, OCF/EV is marked missing before winsorization and ranking. A negative operating-cash-flow numerator remains valid when enterprise value is positive because it conveys economically relevant information.

Momentum Inputs

The model calculates momentum from the adjClose field returned by FMP's Dividend-Adjusted Price Chart API:

12-1 Momentum = adjClose(t-21) ÷ adjClose(t-252) - 1

6-1 Momentum = adjClose(t-21) ÷ adjClose(t-126) - 1

Here, t is the measurement date and the offsets represent trading sessions. Both windows stop approximately one month before t.

The model uses dividend-adjusted prices so the specified return windows incorporate the effect of dividend adjustments. Higher momentum receives a better score. If the required adjClose observations are unavailable, mark the affected momentum input as missing under the published missing-data rules.

Quality Inputs

Return on invested capital is the profitability component:

ROIC = TTM NOPAT ÷ Average Invested Capital

The calculation uses the following components:

Component

Specification

TTM operating income

Sum of operatingIncome from four distinct, consecutive eligible quarterly income statements

TTM pretax income

Sum of incomeBeforeTax from the same four quarters

TTM income-tax expense

Sum of incomeTaxExpense from the same four quarters

Effective tax rate

TTM income-tax expense ÷ TTM pretax income

TTM NOPAT

TTM operating income × (1 - effective tax rate)

Invested capital at quarter q

totalDebt + totalStockholdersEquity - cashAndCashEquivalents

Average invested capital

(Invested capital at q0 + Invested capital at q-4) ÷ 2

The operating-profit and tax fields come from eligible historical quarterly records in the Income Statement API. Invested-capital fields come from eligible historical quarterly balance sheets.

Quarter q0 is the latest eligible fiscal quarter at the measurement cutoff. Quarter q-4 is the corresponding fiscal quarter one year earlier.

The effective tax rate is valid only when TTM pretax income is strictly positive, TTM income-tax expense is non-negative, and the calculated rate falls between 0 and 1. Otherwise, ROIC is marked missing. ROIC is also marked missing when either invested-capital observation is missing or non-finite, or when average invested capital is zero or negative.

Negative TTM operating income remains valid when the tax rate and invested-capital denominator are valid.

Net debt to EBITDA is the balance-sheet component:

Net Debt/EBITDA = (Total Debt - Cash and Cash Equivalents) ÷ TTM EBITDA

TTM EBITDA is the sum of the ebitda field from the same four distinct, consecutive eligible quarterly income statements used for the other TTM flows. Total debt and cash are taken from the q0 balance sheet.

Lower leverage receives a better score. Net debt to EBITDA is calculated only when TTM EBITDA is strictly positive. If EBITDA is zero, negative, missing, or non-finite, the leverage input is marked missing before winsorization and ranking. Negative net debt is permitted when EBITDA is positive because it identifies a net-cash position.

Align Units and Validate Raw Observations

Normalization cannot repair a bad input. Before any ranking, confirm that percentage fields use a consistent scale, monetary values use compatible currencies, prices correspond to the intended share class, and each TTM observation is the latest eligible record as of the measurement timestamp.

A free cash flow yield of 0.06 and a displayed value of 6% represent the same observation, but they differ by a factor of 100 if stored inconsistently. Convert yields and returns to decimals before calculation, then format percentages only for display.

Duplicate symbols and multiple share classes require an explicit policy. A returned peer set can support reasonableness checks, but the factor universe should not change dynamically simply because a peer list changes. Use peer data only where the peer-selection rule itself is versioned.

Save input diagnostics before winsorization, including eligible count, missing count, median, 1st and 99th percentiles, minimum, maximum, and the number of observations altered by each cleaning rule. These checks make silent data shifts easier to detect.

Control Extreme Values Before Ranking

Cross-sectional financial ratios often have long tails. A company with an unusually small denominator can produce a yield or leverage ratio large enough to dominate a z-score.

The illustrative model therefore winsorizes each input at the 5th and 95th percentiles of its normalization group:

Winsorized(x) = min(max(x, Q5,g), Q95,g)

Here, Q5,g and Q95,g are the group's 5th and 95th percentiles.

Winsorization does not delete companies or alter the order of observations within the uncapped range. In this percentile-based model, its central role is to pool observations below the lower cutoff and above the upper cutoff at their respective boundaries. Those pooled observations become ties before percentile ranking.

The cutoff remains a researcher-created choice. A 1st/99th-percentile rule preserves more tail information, while a 10th/90th-percentile rule is more aggressive. Apply one documented rule consistently across rebalances unless a model revision explicitly changes it.

Small sectors create an additional problem because their percentiles can be unstable. This model uses sector-neutral normalization only when a sector has at least 20 eligible observations for an input. Otherwise, it falls back to the full universe and records the fallback flag.

Convert Inputs to Comparable Scores

Percentile ranks provide an intuitive common scale. A normalization group must contain at least two valid observations. When a sector does not meet the minimum group size, the model falls back to the full universe. If the fallback group also contains fewer than two valid observations, the input cannot be scored and is marked missing.

For a higher-is-better input with N valid observations:

Score = 100 × (Average Rank - 1) ÷ (N - 1)

Ranks are assigned in ascending order, and tied observations receive their average rank. Because winsorization can create ties at both tails, the lowest observed score may be greater than 0 and the highest may be less than 100.

A score of 80 means the observation's average rank maps to 80 on this transformed scale. It does not necessarily mean the observation exceeds exactly 80% of the raw group.

For a lower-is-better metric such as net debt to EBITDA:

Score = 100 - Higher-Is-Better Percentile Score

Percentile ranks are robust and easy to explain, but they discard information about the distance between observations. A z-score model preserves more distance information:

Z = (x - Group Mean) ÷ Group Standard Deviation

A z-score model can cap normalized values before combining them. This article uses percentile ranks because interpretability is the priority. Do not mix percentile scores and raw z-scores inside one weighted sum without first placing them on a compatible scale.

Use Sector-Neutral and Cross-Universe Scores Deliberately

Value and quality often differ structurally by sector. Software companies, utilities, industrial businesses, and other non-financial sectors operate with different capital structures, margin profiles, and valuation ranges. Ranking value and quality inputs within a sector reduces the chance that the composite becomes an unintended sector allocation.

Momentum is normalized across the full universe in this example because a price return has the same unit across sectors and the model is intended to capture broad relative strength. That remains a choice, not a universal rule. A sector-neutral momentum score may be appropriate when the research objective is stock selection within industries rather than exposure to sector trends.

Sector-neutral scoring does not remove all industry effects. Broad sectors can contain very different business models. Narrower peer groups can improve comparability, but they also reduce sample size and increase sensitivity to classification changes. A peer-group comparison built from standardized financial ratios illustrates why the comparison set should be deliberate rather than assumed.

Combine Scores Without Hiding the Weights

The composite score is the weighted average of the six normalized inputs:

Composite = 0.20 × FCF Yield Score + 0.15 × OCF/EV Score + 0.20 × 12-1 Momentum Score + 0.15 × 6-1 Momentum Score + 0.20 × ROIC Score + 0.10 × Net Debt/EBITDA Score

This formula keeps the contribution of every input visible. It also makes sensitivity testing straightforward: change one weight, rerun the ranking, and measure how much the order changes.

The model should not add a hidden bonus, analyst override, or qualitative adjustment after the calculation. If another signal matters, it should appear as a defined input or as a separate review field.

Read Raw Inputs, Normalized Scores, and Composite Scores Together

The following hypothetical example shows three companies after universe filtering. The raw values are illustrative and do not represent live FMP data. Percentile scores are simplified for readability and assumed to have been calculated against each company's full eligible normalization group, not only the three displayed rows.

Company

FCF yield

OCF/EV

12-1 return

6-1 return

ROIC

Net debt/EBITDA

Alpha Systems

7.2%

8.1%

18.0%

9.0%

19.0%

0.8x

Beacon Retail

4.5%

5.6%

31.0%

17.0%

13.0%

2.4x

Crest Industrial

9.0%

10.2%

-4.0%

2.0%

16.0%

1.5x

Company

FCF yield score

OCF/EV score

12-1 score

6-1 score

ROIC score

Leverage score

Composite

Alpha Systems

72

68

64

60

81

79

70.5

Beacon Retail

46

51

86

82

57

42

62.0

Crest Industrial

88

84

28

44

70

63

62.7

For Alpha Systems:

Composite = 0.20(72) + 0.15(68) + 0.20(64) + 0.15(60) + 0.20(81) + 0.10(79) = 70.5

Alpha ranks first because it combines above-average results across all three factor families. Beacon's strong momentum is partly offset by weaker value, quality, and leverage scores. Crest's value and quality are stronger, but negative 12-1 momentum reduces its composite.

The example shows the intended trade-off. It does not represent a live ranking, investment recommendation, or FMP-produced score.

Handle Missing Values Explicitly

Missing data should not automatically become zero. A zero score says the company is the weakest valid observation. A missing value says the input could not be measured. Treating them as identical creates an undocumented penalty.

This illustrative model requires at least one valid input in each factor family and at least five of the six inputs overall. When exactly one input is missing, its weight is redistributed proportionally within the same family. If both inputs in a family are missing, the company receives no composite score.

For example, if OCF/EV is missing but FCF yield is valid, the value family's full 35% weight moves to FCF yield. That preserves the intended family weight but increases concentration in one metric. The output should therefore include a coverage count and a missing-input flag so users can distinguish a fully observed score from a partially observed one.

An alternative is to impute the group median, which produces a neutral score near 50. That supports broader coverage but can create the appearance of information where none exists. If imputation is used, disclose the rule and identify every affected observation.

Set Rebalance and Data-Lag Rules

A monthly rebalance is a reasonable starting point for this model. Momentum changes continuously, while TTM fundamentals update around reporting events. Monthly scoring captures price changes without implying that fundamental data requires daily recomputation.

Use the last trading day of each month as the measurement date and form the new ranking only after all eligibility cutoffs are applied. For a live screen, use the latest available TTM record as of the run time. For historical research, use only information that was available under the stated filing and availability rules.

Researchers should also decide how to handle a company that reports between the measurement date and trade date. A defensible process uses information available at the measurement timestamp, then applies the portfolio or watchlist on the next scheduled execution date. Changing that rule after reviewing results introduces discretion that is difficult to reproduce.

Rebalancing more frequently may increase turnover without improving signal quality. A complete evaluation should report rank stability, constituent turnover, transaction-cost assumptions, and performance before and after estimated costs.

Protect Historical Research From Look-Ahead and Survivorship Bias

A current screen and a historical backtest are different data problems. Current TTM endpoints are useful for forming a present-day ranking, but a historical simulation needs records that reflect what was knowable on every earlier measurement date.

Look-ahead bias occurs when the test uses a filing before it was public, uses a restated value as though it were known earlier, or calculates a historical ratio with a later market price.

Survivorship bias occurs when the test applies today's active universe to the past, excluding companies that were delisted, acquired, or failed. The same issue appears when interpreting historical index data: past comparisons need the membership and inputs that existed at the time rather than a modern constituent list.

A robust historical process should preserve:

  • Point-in-time universe membership, including delisted securities where the dataset and research license support them.
  • Filing or acceptance dates, not only fiscal period-end dates.
  • The price series and corporate-action convention used at each measurement date.
  • The model version, universe rules, winsorization boundaries, group membership, and missing-data decisions for every rebalance.
  • A realistic delay between data availability, signal calculation, and implementation.

Sector and industry classifications can also change. Using today's classification for a past ranking may alter both the peer group and the normalized score. Where historical classifications are unavailable, disclose the limitation and test its sensitivity.

Build the Data Pipeline in a Reproducible Order

A practical implementation can follow this sequence:

  1. Freeze the measurement cutoff at June 30, 2026, 4:30 p.m. America/New_York, and load a dated security-universe snapshot containing securities active at that cutoff.
  2. Apply the security-type and Financial Services exclusions. Apply the $1 billion threshold using the Historical Market Capitalization API and calculate the liquidity threshold from the 20 historical trading sessions ending on June 30.
  3. Retrieve historical quarterly income statements, balance sheets, and cash-flow statements. Apply the stated acceptedDate and filingDate availability rules without assuming that an unzoned timestamp uses Eastern Time.
  4. Select four distinct, consecutive eligible fiscal quarters for each TTM flow. Do not combine annual and quarterly records or count duplicate versions of the same quarter.
  5. Construct TTM free cash flow, operating cash flow, operating income, pretax income, income-tax expense, and EBITDA. Retrieve the q0 and q-4 balance sheets required for average invested capital.
  6. Calculate measurement-date enterprise value, the two value inputs, ROIC, and net debt to EBITDA using the published definitions and invalid-denominator rules.
  7. Retrieve dividend-adjusted momentum prices and calculate both momentum inputs using adjClose and the stated trading-session offsets.
  8. Validate units, dates, currencies, duplicate securities, statement continuity, invalid denominators, and missing values before normalization.
  9. Winsorize within the published groups, assign average-rank percentile scores, invert the lower-is-better leverage input, and apply the six published weights.
  10. Save the dated universe, raw source records, eligibility decisions, cleaned inputs, winsorization boundaries, group membership, normalized scores, composite scores, coverage flags, source timestamps, and model version as separate audit layers.

A live screen is a separate implementation. It may begin with current screener, profile, ratio, and key-metric data, but those current records should not be inserted into the June 30 historical reconstruction. The endpoint mapping, validation, and ranking logic required for a current TTM screen are addressed in Beyond P/E: Screening, Scoring, and Ranking Stocks With FMP TTM Ratios API.

Test Whether the Model Is Doing What It Claims

Before evaluating returns, test the score itself.

  1. Check the distribution of each input and normalized score, the correlation between inputs, the concentration of top-ranked names by sector, and the proportion of partially observed companies. A composite called diversified may still be dominated by one factor if the inputs are highly correlated.
  2. Change winsorization from the 5th/95th to the 1st/99th percentiles, compare sector-neutral with cross-universe ranks, equal-weight the six inputs, and remove one input at a time. If small specification changes completely reorder the top names, the model is fragile.
  3. Evaluate more than headline return. Useful diagnostics include rank monotonicity across score buckets, turnover, drawdown, sector exposure, size exposure, liquidity, transaction costs, and performance across different market regimes.
  4. Treat the final output as a research ranking, not an investment recommendation. A high score identifies companies that rank well under the published definitions at one measurement date. It does not account for every risk, catalyst, portfolio constraint, or qualitative consideration.

What a Transparent Factor Output Should Show

At minimum, the published or internal output should include the measurement date, universe definition, factor version, composite score, family scores, raw inputs, normalized inputs, sector or industry group, missing-data count, and any fallback flags. Users should also be able to retrieve the formula and weight attached to each field.

That level of detail makes disagreement useful. A reviewer can challenge sector-neutral value, the exclusion of the latest month from momentum, or the treatment of negative EBITDA without questioning where the result came from. Transparent construction does not remove judgment. It makes judgment visible.

Once the specification is fixed, create a free FMP API key to pull the market and financial inputs needed to test the model's own definitions rather than relying on a composite score alone.

FAQs

What is the difference between a factor and a screening rule?

A screening rule applies a threshold, such as market capitalization above $1 billion or free cash flow yield above 5%. A factor ranks eligible observations along a continuous scale. Screens determine which securities enter the research set; factors compare the securities that remain.

Why use percentile ranks instead of raw ratios?

Raw ratios have different units and ranges, so they cannot be added meaningfully without transformation. Percentile ranks place every input on a common bounded scale and limit the influence of extreme distances. Tied observations receive average ranks, so the scores observed in a particular group do not necessarily reach exactly 0 and 100.

Should value and quality always be sector-neutral?

No. Sector-neutral scoring is useful when structural industry differences would otherwise dominate the ranking. Cross-universe scoring is appropriate when those differences are part of the intended exposure. The choice should match the research objective and remain consistent across rebalances.

How should negative earnings or EBITDA be handled?

The rule depends on the formula. Negative cash-flow yields can remain valid observations and be winsorized. Ratios with an economically invalid denominator, such as net debt divided by negative EBITDA under this specification, should be marked missing or addressed through a separately defined distress signal rather than forced into the same ranking.

Can current TTM data be used for a historical backtest?

Not by itself. A historical test needs point-in-time fundamentals and the dates on which those values became public. Using the latest TTM value at an earlier measurement date creates look-ahead bias, even if the fiscal period overlaps the historical test window.

Does a high composite score mean a stock is a buy?

No. It means the security ranks well under this illustrative model's disclosed definitions, normalization rules, and weights at one measurement date. The score is a research tool, not an FMP rating, recommendation, price target, or substitute for broader due diligence and portfolio-risk analysis.

About the Author

Sanzhi Kobzhan
Sanzhi Kobzhan

Treasury, trading, liquidity, and equity analysis for investors

Sanzhi writes for FMP with a focus on equity analysis, valuation, market data, and practical investment decision-making. He has worked across financial institutions in treasury, trading, and liquidity roles, bringing hands-on experience in investment analysis, market execution, risk, and strategy. His work focuses on helping readers interpret financial data with clarity, discipline, and an institutional market perspective.

Financial data for every need

Real-time quotes and 30+ years of historical data, including prices, fundamentals, and insider transactions — all accessible via API.