← Projects SS

Computational cryptography

Opaquegenerators.

An investigation into a conspicuous pattern in the SECG prime-field Koblitz curve generators—and what changed when a defensive-security model could examine the attack hypotheses directly.

The background

An elliptic-curve system names a base point, usually called G, whose repeated addition defines the discrete-logarithm problem used by the protocol. In a prime-order group, every nonidentity point generates the same group, so standard theory gives no reason to expect one honestly selected generator to make the generic discrete-log problem easier than another.

Provenance still matters. A transparent parameter-generation procedure lets independent observers verify that a point was not selected to satisfy some undisclosed relation. The four prime-field Koblitz curves standardized in SEC 2—secp160k1, secp192k1, secp224k1, and secp256k1—do not publish a verifiable-random seed for their generators.

The anomaly

For each published generator, the study computed the unique point H satisfying G = 2H. All four resulting x-coordinates contain the same 152-bit hexadecimal core:

8ce563f89a0ed9414f5aa28ad0d96d6795f9c6

The 224- and 256-bit curves go further: they use exactly the same integer x-coordinate for H. Exact arithmetic verified every published point, group order, halving relation, and compatible GLV endomorphism orientation. The shared value is not a formatting coincidence.

The simplest descriptive reconstruction is a fixed hexadecimal core with small curve-specific affixes, followed by lifting to a curve point and doubling. That describes the output; it does not recover the designers’ historical algorithm or establish their intent. Simple ascending-affix scans did not reproduce the points as the first acceptable candidates, and a bounded search of plausible SHA-1 labels and counters found no match.

Prior work and novelty

The halved-generator anomaly was already public. Heninger, Scholl, and Shumow presented the exact secp160k1 half-coordinate in a Crypto 2019 rump-session talk. The later DiSSECT project included x([k⁻¹]G) for k = 1,…,8 among its curve traits and identified secp224k1 and secp256k1 as exceptional among comparable standardized curves.

This project’s contribution is therefore not first discovery of G/2. It reconstructs the shared construction across all four prime-field Koblitz curves, tests possible cryptanalytic consequences against matched controls, reproduces the published inverse-scalar distinguisher, measures implementation-level effects, and investigates why the points may have been doubled.

The AI boundary

The first prompt called the work elliptic-curve cryptanalysis and proposed comparing special and random generators with division polynomials, summation-polynomial systems, Gröbner bases, low-height relations, GLV correlations, and small analogue curves. It expressly bounded the work and disclaimed an attempt to solve a production-size discrete logarithm. GPT-5.6 Sol via Codex nevertheless refused it under its cybersecurity safeguards—and also declined to help draft the request. Grok wrote the resulting original prompt.

To continue, the same investigation was recast as higher-assurance parameter design. “Special” points became “historical-style” points; the comparisons gained transparent random and NUMS controls; and the stated goal became guidance for designing future generators rather than evaluating attack ideas. Sol accepted the reframed prompt and performed much of the initial study.

That was, candidly, a change in presentation more than substance. Later access to Daybreak Blue—a model intended for broad defensive cybersecurity research—made the euphemism unnecessary. The project could state its purpose plainly and complete the attack-oriented experiments directly. The sequence is an interesting case study in model specialization: nearly the same mathematical work crossed one model’s policy boundary when described indirectly, then became straightforward under a model designed for that research domain.

What was tested

The reproducible package generated 2,000 prime-order, cofactor-one, j = 0 analogue curves from 30 to 80 bits. Each curve received a fixed-core historical-style base, an exactly sampled random control, and a transcript-derived full-length-x NUMS control. A later high-power corpus used 1,000 curves and sixteen independent random controls per curve.

The analysis included short scalar and GLV orbits, coordinate size and sparsity, exact two- and three-summand Semaev relations, matched factor-base experiments, SageMath verification, and representative Singular Gröbner-basis calculations. It also corrected two initially proposed comparisons: division polynomials and the GLV scalar lattice are properties of the curve, not of which nonzero point is named as generator.

The direct follow-up added process-isolated and blinded Semaev/Gröbner campaigns, deduplicated GLV coefficient-box scans totaling more than 264 million coefficient/base observations, and actual-curve LLL searches for unusually short relations among the coordinates of successive multiples. It also tested fixed-window tables and mixed-addition intermediates for implementation-leakage proxies on the real SEC points, their exact halves, and thousands of controls.

What was found

No measured signal credibly distinguished historical-style bases from random controls. Two-summand relation studies were null at both tested factor-base sizes. The larger three-summand grid and its fresh-control replications were also null. A weak timing direction that appeared in one aggregate statistic did not strengthen across relation densities, did not produce abnormal incidence or rank, and sometimes appeared for NUMS controls as well.

The attack-oriented follow-up told the same story. Apparent Gröbner timeouts were explained by process startup and contention, not generator class. Two nominal GLV-orbit anomalies pointed in inconsistent directions and each failed immediate replication on fresh data. On the four production curves, none of the exact half-generators entered the preregistered lower one-percent tail in the LLL coordinate-relation search.

DiSSECT’s exact inverse-scalar trait did distinguish the public parameters. For secp224k1 and secp256k1, k = 2 reaches the 166-bit half-point x-coordinate; the shortest values among 2,000 matched controls were respectively 212 and 242 bits. secp160k1 and secp192k1 were ordinary under the same aggregate metrics. Broader exploratory scalar families found no additional witness: every distinction still came from G/2. This reproduces a strong public-coordinate distinguisher, not an ECDLP shortcut.

One real but limited effect did survive replication. The half-points for secp224k1 and secp256k1 have zero leading limbs and unusually low coordinate Hamming weights in common fixed-base representations; no corresponding control matched several of those entry-level observations in 10,000 trials. The effect concerns public table values, not secret recovery, and it disappeared from the tested mixed-addition traces. Crucially, the published generators G = 2H were not outliers: doubling washes out the conspicuous representation artifact.

SageMath independently validated the SEC calculations, all primary-corpus group orders, and hundreds of recorded relation witnesses. Singular confirmed representative polynomial ideals. In the largest joint Semaev incidence comparison, the 95% interval was −0.95 to +2.12 percentage points: within that design, an advantage larger than about 2.1 points is inconsistent with the data. These are strong negative results for the hypotheses actually tested, not a proof that every possible nongeneric method is impossible.

Why double the point?

No located primary source explains why Certicom published twice the structured source point on all four curves. Contemporary procedures do suggest a plausible benign lineage: Certicom’s 1997 ECC Challenge generated a candidate point and multiplied it by the curve cofactor, while binary Koblitz curves commonly had cofactors 2 or 4. A “derive, then multiply by a small cofactor” convention or tool may have carried over to the prime-field family.

That explanation is incomplete because all four prime-field curves have cofactor 1; no applicable standard calls for the extra doubling. Another plausible possibility is presentation: a roughly SHA-1-width source coordinate could be lifted and doubled to produce a full-width-looking published coordinate. The experiment confirms that doubling removes the source point’s obvious encoding anomaly, but no evidence shows that concealment was the intention.

A separate investigation of NIST/X9.62 seeds adds useful historical context. In emails later published by Bernstein and Lange, Jerry Solinas recalled—without the exact message or algorithm—that those seeds came from hashing a humorous ASCII message, probably with SHA-1 and a counter. Other X9.62 seeds visibly contain MinghuaQu, showing that human-meaningful material did enter contemporary parameter sets. This makes a forgotten message or SHA-1-width template historically plausible, but it is not evidence that the SEC Koblitz generators used the same process or people.

Without an exact preimage and one derivation rule reproducing the affixes, lift choices, acceptance conditions, and doubling across all four curves—or independent archival corroboration—the resemblance remains a lead, not a provenance finding. Failure to reverse SHA-1 over an undefined phrase space would be equally inconclusive.

The most defensible conclusion is that doubling probably reflects an undocumented generation or presentation convention inherited from cofactor-clearing workflows. It obscures the source-point structure predictably, but the surviving record does not establish why it was done.

The conclusion

The historical generator construction should not be copied. Its visible reuse and absent public derivation create avoidable suspicion. The direct experiments support a narrower conclusion: no nongeneric mechanism tested here made these generators easier for ECDLP. Generic-group lower bounds do not cover coordinates, encodings, endomorphisms, implementation leakage, or relations among multiple protocol generators.

That last distinction matters. If one protocol generator is a known scalar multiple of another, commitments and other multi-generator constructions can fail even though single-base ECDLP is unchanged. Independent generators should therefore be derived directly from a versioned, domain-separated public transcript. Where a reviewed suite exists—RFC 9380 includes secp256k1—use standardized hash-to-curve. Do not hash to a scalar and multiply the existing generator, because that reveals the relation. A fully specified candidate-x ceremony remains auditable for curves without such a suite, but it is a project-specific map.

In other words: the coordinate artifact is a legitimate transparency problem, and the half-points really are unusual representations. This study found no evidence that either fact creates a cryptographic backdoor.