//nefariousplan

CVE-2026-7669: SGLang Treated trust_remote_code=False As a Suggestion

pattern

cve

proof of concept

In SGLang 0.5.10 through current main, the operator who passes trust_remote_code=False receives the same tokenizer as the operator who passes trust_remote_code=True for any model carrying a custom tokenizer_class. Twelve lines in get_tokenizer observe that transformers v5 returned its generic TokenizersBackend fallback and re-invoke AutoTokenizer.from_pretrained with trust_remote_code=True. The introducing PR was titled "Upgrade transformers==5.3.0" and shipped the override with a logger.info notice. The PR that removed the notice five days later was titled "docs: improve CI and testing documentation."

The override is twelve lines, and the comment above them names the inversion

The block as introduced in commit d1e95af28 (PR #17784, "Upgrade transformers==5.3.0", merged 2026-03-18):

# Transformers v5 may silently fall back to a generic TokenizersBackend
# when trust_remote_code=False and the model requires a custom tokenizer.
# Detect this and auto-retry with trust_remote_code=True.
if not trust_remote_code and type(tokenizer).__name__ == "TokenizersBackend":
    logger.info(
        "Detected generic TokenizersBackend for %s, "
        "retrying with trust_remote_code=True",
        tokenizer_name,
    )
    tokenizer = AutoTokenizer.from_pretrained(
        tokenizer_name,
        *args,
        trust_remote_code=True,
        tokenizer_revision=tokenizer_revision,
        clean_up_tokenization_spaces=False,
        **kwargs,
    )

The comment is testimony. Line one describes transformers v5's behavior as a "silent fall back." Line three describes the SGLang response as an "auto-retry with trust_remote_code=True." Whatever the author meant by criticizing transformers for falling back silently in line one, the action three lines down is the same operation, with the security flag flipped. The author named the inversion. They documented it. They shipped it.

What the comment does not name is what trust_remote_code=False was for. The HuggingFace documentation is unambiguous: trust_remote_code is the boolean that gates the execution of tokenizer.py (and configuration.py, and modeling.py) from a model repository at load time. The operator who passes False is asserting that no repository on the Hub gets to execute Python in their inference process. SGLang's wrapper observed that this assertion produced a generic class for one specific shape of model, decided the generic class was the wrong answer, and called the underlying loader again on the operator's behalf with the assertion inverted.

The type(tokenizer).__name__ == "TokenizersBackend" predicate is the only gate between this inversion and a no-op. TokenizersBackend is the class transformers v5 returns when a model's declared model_type has no entry in TOKENIZER_MAPPING_NAMES. The PoC's trigger model declares model_type: "gpt2" in config.json (which is in the registry) and a custom tokenizer_class in tokenizer_config.json (which is what causes transformers to use the fallback class). Both files are attacker-controlled. Both are part of the standard HuggingFace Hub layout. Nothing in the trigger condition requires anything but a free Hub upload.

The notice was audible for five days

PR #17784, merged 2026-03-18, was titled "Upgrade transformers==5.3.0." It is a real version-bump PR. It touches the rotary embedding factory, get_hf_text_config, the model loader, the rope-parameters compatibility helpers, the VLM encoder modules. Among the legitimate v5 compatibility work, it adds the twelve lines above to get_tokenizer, logger.info call included.

PR #21202, merged 2026-03-23, was titled "docs: improve CI and testing documentation." It touches 119 files. 519 insertions. 809 deletions. Most of the change is real: README updates, CI workflow rewrites, deletion of test/srt/experiment_runner.py and three sibling cleanup scripts, formatting on JIT kernel test files. Five of the deletions are in python/sglang/srt/utils/hf_transformers_utils.py:

     if not trust_remote_code and type(tokenizer).__name__ == "TokenizersBackend":
-        logger.info(
-            "Detected generic TokenizersBackend for %s, "
-            "retrying with trust_remote_code=True",
-            tokenizer_name,
-        )
         tokenizer = AutoTokenizer.from_pretrained(
             tokenizer_name,
             *args,
             trust_remote_code=True,

The PR's commit message does not mention the change. The PR's title does not mention the change. The diff classifies the deletion alongside the deletion of experiment-runner scripts and the rewriting of test workflows. The change passed review under a title that read as cleanup.

After 2026-03-23, the override emits nothing. The Python logging module's root logger captures nothing. The sglang logger captures nothing. The transformers logger sees only the second AutoTokenizer.from_pretrained call, which on the wire is indistinguishable from a legitimate launch under --trust-remote-code. The PoC's PHASE-2-silent claim instruments DEBUG capture on both the root logger and the sglang logger across the full call and asserts no log line mentions trust_remote_code. The assertion passes.

The commit that introduced the override named only a version bump. The commit that silenced it named only a docs cleanup.

The reachability surface is server launch, not runtime

get_tokenizer has four reachable callers in the SGLang server:

python/sglang/srt/managers/tokenizer_manager.py:323     TokenizerManager.__init__
python/sglang/srt/managers/scheduler.py:585             Scheduler.__init__
python/sglang/srt/managers/tp_worker.py:282             TPWorker.__init__
python/sglang/srt/managers/detokenizer_manager.py:108   DetokenizerManager.__init__

All four are constructors. They run once per fresh process, at server launch, against the model path the operator supplied on the command line. The PoC's CHAIN-1 claim grep'd python/sglang/srt/entrypoints/http_server.py for references to get_tokenizer and matched zero handlers. No HTTP endpoint reaches the override.

This matters for the disclosure framing. SGLang has a real default-no-auth bypass: when api_key=None and admin_api_key=None, all ADMIN_OPTIONAL endpoints are reachable without credentials. The PoC's AUTH-1, AUTH-2, and AUTH-3 claims verify the bypass against the live decide_request_auth primitive. None of those endpoints reach get_tokenizer. The auth bypass and the trust-remote-code override are separate bugs that do not chain. The override fires once per launch from the operator's --model-path argument, not from any post-startup HTTP call.

What this means in practice: an attacker who can persuade an operator to launch SGLang against a specific Hub model achieves arbitrary Python execution inside the SGLang process at startup. The persuasion vector is the open Hub. The operator did not need to disable any security control. They explicitly enabled the security control by passing trust_remote_code=False. The wrapper layer overrode them.

The pattern is setting-was-advisory, and the substrate is wrapper layers

The shape has a name. setting-was-advisory is the class where a caller passes a security-relevant flag with the API surface of an enforced control, and the wrapper layer treats the flag as a hint. The flag's contract says False means False. The wrapper's behavior is that False means False unless the framework's safe response to False is inconvenient, in which case True. The caller has no path to learn that the substitution happened.

The pattern is adjacent to fail-open-intercept and not the same as it. In Tomcat's EncryptInterceptor, the security gate could not decrypt the message and forwarded the bytes anyway. Here, the gate did make a verdict. Transformers v5 returned TokenizersBackend precisely because the caller said trust_remote_code=False. That was the correct response. The SGLang wrapper inspected the correct response, decided it was inconvenient, and called the loader again with the security flag inverted to get a more convenient one.

The substrate where this pattern grows is wrapper layers around third-party libraries. The library upgrades and changes a return shape. The wrapper's downstream code expects a tokenizer with model_max_length and other model-specific attributes that the generic TokenizersBackend does not carry. The wrapper author's options when they observe the new return shape are three:

  1. Accept TokenizersBackend and let downstream code fail loudly when it tries to access an absent attribute. The cost is a regression on models that worked under transformers v4.
  2. Raise from get_tokenizer and instruct the operator to pass --trust-remote-code if they need the custom tokenizer. The cost is operator-facing breakage on existing deployments.
  3. Silently retry with trust_remote_code=True, absorb the framework's change inside the wrapper, and document the inversion in a comment. The cost is the security flag inversion this CVE names.

The wrapper author chose option three. They wrote a comment naming the inversion. They added a logger.info to announce it at runtime. Five days later the announcement was removed. The wrapper exists, in production, to absorb a transformers v5 API change at the cost of silently violating the caller's security assertion. The patch this CVE asks for is a five-line deletion. The discipline the substrate asks for is harder: a wrapper that responds to "the framework's safe default does not carry the attributes I want" by retrying with safety off is the same shape, regardless of language. The class will keep producing CVEs of this form until wrapper authors agree that re-invoking with a flipped security flag is a different operation that requires the caller's consent, not the wrapper's discretion.

The disclosure timeline is one-sided

Date Event
2026-03-18 PR #17784 merged. Override introduced with logger.info notice.
2026-03-23 PR #21202 merged. logger.info deleted under the title "docs: improve CI and testing documentation."
2026-04-07 Nick Gould discovers the override, builds a working PoC, reports via SGLang Private Vulnerability Reporting and via VulDB.
2026-05-03 CVE-2026-7669 assigned by VulDB.
2026-05-04 PoC published.

NVD's record for CVE-2026-7669 reads, in part, "The vendor was contacted early about this disclosure but did not respond in any way." The CVE was assigned by an external coordinator on the strength of the reporter's submission and the working PoC, not by the vendor. As of this writing, SGLang's main branch on GitHub still contains the override block. The patch the PoC's PHASE-1b verifies is surgical: delete the twelve lines. The vendor has not deleted them.

The HEAD of main has refactored the function into a new module, python/sglang/srt/utils/hf_transformers/tokenizer.py, where the equivalent function _resolve_tokenizers_backend logs a warning on every retry and only attempts dynamic-module loading when the caller already passed trust_remote_code=True. That refactor is a real improvement, and it does not appear in any released version. The pinned-version PoC continues to fire because the override block, as introduced in March 2026 and as silenced five days later, is what pip install sglang==0.5.10 and every pip-installable release through current still ship.

PoC: gouldnicholas/CVE-2026-7669-PoC

The override is still there. The log is not.