A coefficient is not a finding until a reader can see why that number is not just selection, reverse causality, or an omitted variable wearing a t-statistic. Identification is the written argument that converts a regression into a claim. In Indian and other emerging-economy manuscripts, this argument is often missing, or it is replaced by a sentence that says “we use firm and year fixed effects.” Reviewer 2 is not impressed. This briefing shows how to write the identification section so the design and the verbs stay in the same paper.
The audience is the same as the rest of the Research Series: project writers, master’s researchers, doctoral candidates, post-docs and early faculty. The examples lean toward finance, accounting, ESG and firm-level work, but the writing discipline travels.
1. What identification actually is
Identification is not a method brand. Difference-in-differences, instrumental variables, regression discontinuity, matching and staggered adoption are tools. Identification is the answer to a narrower question: given the way units entered treatment, what variation are you using, and why is that variation not the same thing as the outcome’s own determinants?
Write that answer in prose before you name the estimator. If you cannot narrate the source of variation in two sentences, you do not yet have an identification strategy. You have a specification.
Three failures dominate early-career drafts from the region. First, the paper claims “impact” or “effect” after a pooled OLS with industry dummies. Second, the paper lists robustness tests as if quantity substitutes for a threat model. Third, the paper treats an Indian policy circular as an automatic experiment without showing who was actually bound by the circular and who could opt out.
2. Match verbs to design strength
Editors and reviewers read verbs before they read standard errors. If the design is associational, the verbs must stay associational. Inflated language is not ambition; it is a credibility leak.
| Design you actually have | Verbs you may use | Verbs to retire |
|---|---|---|
| Cross-section with observables | associated with; co-moves; predicts | causes; impact of; the effect of |
| Panel with unit and time FE | within-unit change; conditional association | causal effect (unless threats are closed) |
| Credible shock / DID / RD / IV | effect; response to; estimated impact | “proves”; “establishes once and for all” |
| Qualitative process evidence | how actors interpret; sequence of moves | average treatment effect language |
The identification section is where you justify the verb choice in public. If you later write “the effect of board independence” in the abstract, the identification section must already have earned that noun.
3. Anatomy of a review-proof identification section
Aim for 400–800 words in a typical empirical article, placed after data construction and before or immediately after the baseline specification. Structure it as six short blocks. Do not bury it inside a methods paragraph that also discusses winsorisation.
3.1 The target parameter
State what you want to learn in one sentence. “We want the change in capital expenditure that would occur if a firm moved from unassured to assured sustainability disclosure, holding time-invariant firm traits and common year shocks fixed.” If you cannot name the target parameter, you will estimate whatever falls out of the software.
3.2 The source of variation
Name the comparison. Who switches? Who stays put? Is the switch staggered? Is it a threshold? Is it a bank-driven reallocation? Readers should be able to draw the comparison groups on a napkin.
3.3 Why that variation is not the outcome in disguise
This is the identifying assumption in ordinary language. Parallel trends, exclusion, continuity at the cutoff, conditional independence after observables—pick the one that matches the design and say what it would take for the assumption to fail in your setting.
3.4 The most dangerous alternative story
Write the Reviewer 2 paragraph yourself. “A plausible alternative is that high-growth firms both adopt assurance and raise investment.” Then say what you do to that story: pre-trends, placebo outcomes, timing tests, a bounding exercise, a selection-into-treatment model. One alternative story treated seriously is worth five generic robustness tables.
3.5 What the design cannot answer
Honesty is identification. If you cannot separate price from quantity, say so. If you cannot speak to informal firms outside the CMIE or Prowess universe, say so. Reviewers punish hidden limits more than stated limits.
3.6 Mapping to tables
End with a one-line map: Table 3 is the baseline comparison; Table 4 is the pre-trend and placebo set; Table 5 is the heterogeneity that tests the mechanism implied by the identifying story. The identification section should make the table order feel inevitable.
4. Using Indian institutions as design, not colour
Emerging-economy settings are often sold as novelty. They are more useful as design. Indian institutions produce discontinuities, staggered mandates, eligibility thresholds and enforcement gradients that richer-country panels sometimes lack. The identification section must show that you used the institution as a source of variation, not as a postcard.
Useful institutional features, when documented carefully, include reporting thresholds (BRSR applicability, listing status, paid-up capital or turnover cutoffs), credit-policy boundaries (priority-sector classification, targeted refinance windows), board and audit rules that bind listed firms differently from unlisted peers, and state-level implementation of central schemes. None of these is automatically an experiment. Each requires a paragraph on who is bound, who can delay, and who can reclassify.
Treat as design
Binding threshold; staggered compliance dates; enforcement intensity that varies by regulator or exchange; eligibility that is not chosen by the outcome variable.Treat as colour
“India is under-studied”; “unique cultural context”; “growing economy”; any sentence that could be copied into a tourism brochure.Document the rule as an empiricist, not as a commentator. Cite the circular, the effective date, the grandfathering clause, and the exemptions. If promoters can restructure to fall below a threshold, say so and show whether they did. Identification dies in the exemption footnote that the author never read.
Further reading on DrDPKlass: How to Write a Contribution Paragraph Editors Cannot Ignore and From Desk Rejection to Acceptance.
5. Name the threats before Reviewer 2 does
A professional identification section is a threat inventory with a disposition next to each threat. Use a short explicit list. Reviewers should see that you anticipated them.
Selection into treatment. Firms that adopt a practice are not random. Show observables at adoption, a matching or entropy-balance check if you must stay associational, and a discussion of unobservables that would still remain.
Reverse causality. Outcomes can cause the “treatment” label. Timing plots, lag structures and institutional calendars help only if the calendar is real. A one-year lag is not identification if the firm decided both variables in the same board meeting.
Omitted time-varying confounders. Unit fixed effects kill only the time-invariant slice. Write down the time-varying story you fear—commodity prices for resource firms, state elections, monsoon, GST transition, pandemic relief—and say which controls or sample splits address it.
Measurement of treatment. In ESG and disclosure work this is often the real identification failure. A dummy for “has a sustainability report” is not the same object as assured, comparable, or decision-relevant disclosure. If treatment is noisy, the coefficient is not a clean effect even with a perfect shock.
Interference and spillovers. Peer firms in the same industry or credit market may respond to a treated neighbour. If you assume SUTVA by silence, a careful reviewer will not.
Staggered timing artefacts. If adoption is staggered, say how you avoid forbidden comparisons of late and early treated units. Naming the estimator (and why you chose it) belongs here, not in an appendix footnote.
Reviewer 2 does not require a perfect design. Reviewer 2 requires evidence that you know which imperfections you are asking the reader to live with.
6. When you have no instrument and no shock
Most Indian doctoral datasets will not contain a clean natural experiment. That is not a reason to invent one, and it is not a reason to pretend OLS is IV. It is a reason to write a more modest, more useful paper.
First, downgrade the target parameter in public. “We estimate conditional associations that are informative if the listed confounders are the main threats” is a respectable sentence. Many applied journals will still take the paper if the measurement is new, the mechanism tests are sharp, and the verbs stay honest.
Second, use design-adjacent tools without theatrical claims: within-firm changes, Oster-style bounding, placebo outcomes that should not move, falsification samples (unaffected industries, pre-period leads), and split-sample tests implied by theory. These do not create causality. They shrink the set of stories that can still explain the coefficient.
Third, add process evidence where it earns its keep. Interviews, annual-report language, or implementation timelines can show that the treatment was real for managers, even if the econometrics cannot isolate an average treatment effect. Keep this evidence in the identification logic, not as decoration in a discussion section that nobody maps back to the tables.
Fourth, do not shop for an instrument after seeing the result. A weak or implausible instrument is worse than no instrument. Reviewers in finance and accounting have seen the “rainfall in another state” genre. If the exclusion restriction cannot be defended in a paragraph a non-specialist understands, drop the IV.
Fifth, write a pre-analysis note to yourself even if the journal will never see it. List the comparison groups, the primary outcome, the two main threats, and the verb you are allowed to use. When a new robustness idea appears at 1 a.m., check it against that note. Identification sections collapse when the author keeps adding tests that imply a different target parameter from the one in the introduction.
7. Where the section lives and how it talks to tables
Put the identification argument where a reviewer looking for it will find it: a dedicated subsection with that word in the heading. Do not scatter identifying assumptions across data, methods and robustness. Scattering is how contradictions appear—one paragraph assumes parallel trends, another admits that treated firms were already diverging.
Align three texts: abstract verbs, introduction contribution, identification section. If the introduction promises a shock-based effect, the identification section must describe that shock with dates and bound units. If the identification section is associational, the introduction cannot smuggle causality back in through “implications for policy impact.”
Tables should be labelled as tests of the identifying story, not as a museum of specifications. A robustness table that does not map to a named threat is noise. A short threat-to-table crosswalk in the identification section saves reviewers from hunting and saves you from adding a fifteenth column that tests nothing.
Finally, write the limitations paragraph as a continuation of identification, not as a confession dump. “We cannot speak to unlisted firms” is useful if your mechanism depends on listing-status enforcement. “Future research can use better data” is not identification; it is a placeholder.
8. Field checklist and closing brief
Before you submit, print the identification subsection and tick these items in the margin.
- Target parameter named in one sentence.
- Source of variation named: who changes, who does not, on what calendar.
- Identifying assumption stated in ordinary language, with a concrete failure mode.
- One primary alternative story treated in depth.
- Verbs in abstract, introduction and conclusion match the design.
- Indian (or other local) institution documented as a rule with exemptions, not as ambience.
- Treatment measurement defended; dummy variables justified or replaced.
- Staggered or clustered designs use an estimator that matches the timing.
- Each main robustness table maps to a named threat.
- What the design cannot answer is explicit.
Identification writing is not a performance of econometric fashion. It is courtesy to the reader: you show the comparison you want them to trust, the comparison you cannot defend, and the language that should follow from that distinction. Papers from emerging-economy datasets do not need to apologise for missing instruments. They need to stop borrowing causal verbs they have not paid for.
If you do only one thing after this briefing, rewrite the identification subsection before you touch another robustness column. The extra specification will not save a missing threat model. A clear threat model will often make half the extra specifications unnecessary.