Repository navigation
UN-4224 [FIX] Stop sending temperature to GPT-6 models on Bedrock, OpenAI and Azure OpenAI - #2310
Conversation
…enAI and Azure OpenAI GPT-6 models (luna, sol, astra, 6.1-sol) reject the `temperature` parameter. Our LLM adapters always pass one (the pydantic default, or 1 when reasoning is on), and LiteLLM forwards it on the Bedrock Mantle route, so every GPT-6 Luna call on Mantle failed with "temperature not permitted for this model". GPT-5.6 was never affected because LiteLLM has a GPT-5 rule that drops a non-1 temperature; that rule matches `gpt-5` names only. The Converse route (`us.`/`global.` profile ids) already drops it. - Add a `gpt-6` stem to `_SAMPLING_DEPRECATED_MODEL_STEMS`, the strip already used for Claude Opus 4.7+. The Bedrock, Azure AI Foundry, Vertex and Anthropic adapters call it on their return path. The trailing-edge anchor keeps `gpt-5.6-*` (normalised to `gpt-5-6-*`), `gpt-60` and `gpt-oss-*` out. - Call the strip from the native OpenAI adapter too. - Azure OpenAI: the deployment name need not name the model, so detect GPT-6 from the optional Model field as well. `LLM` sets `cost_model` aside and re-validates without it, where Azure's default temperature of 1 would return, so pin `temperature` to None instead of popping it (LiteLLM omits a None temperature) and stop reasoning from forcing 1. LiteLLM is intentionally not upgraded here: the first release whose bundled registry carries the GPT-6 Mantle entries (1.104.0) needs boto3>=1.43.1, which conflicts with the boto3 1.34 / s3fs 2024.x pins. See UN-4224 for the routing gap that leaves on air-gapped deployments. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
…e per call Review (Greptile P1): for an Azure OpenAI deployment whose name hides GPT-6, the only record of the model that survived `LLM`'s re-validation was the pinned `temperature: None`. `complete()` merges per-call kwargs over the stored ones, so a caller passing `temperature=` (the cloud agentic_table / agentic_extraction workers do) overwrote that marker and the temperature reached a model that rejects it. Detect from the real model id instead of a caller-writable marker. `LLM._revalidate` now feeds the `cost_model` it set aside back into re-validation, after the per-call kwargs so they cannot displace it, and the four completion paths use it. Azure prefers that `cost_model` as the original model id, so detection holds on every pass and the sampling params are simply popped like every other adapter -- the `None` pin and its sticky-marker check are gone. Other adapters are unaffected: none declares `cost_model`, Pydantic ignores unknown keys, and every call site already pops `cost_model` before LiteLLM sees the kwargs. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ge stack LiteLLM 1.104.0 is the first release whose bundled model registry carries the GPT-6 Bedrock Mantle entries, so air-gapped deployments (which fall back to the bundled registry) now route `openai.gpt-6-luna` to Mantle and price GPT-6 calls instead of recording $0. LiteLLM >= 1.100 hard-requires boto3 >= 1.43.1, which drags the storage stack with it; these move together: - litellm 1.96.2 -> 1.104.0 - boto3 / botocore 1.34.x -> 1.43.106, pinned exactly in backend, workers and sdk1: it must stay inside the botocore range aiobotocore 3.9.2 accepts (1.43.101-1.43.106) - s3fs / fsspec 2024.10.0 -> 2026.9.0 (s3fs drops its `[boto3]` extra) - gcsfs 2024.10.0 -> 2026.10.0, adlfs 2024.7 -> 2026.8.0 - google-cloud-storage 2.9.0 -> 3.16.0: every gcsfs compatible with fsspec 2026.x needs google-cloud-storage 3.x. Our direct use (Client, bucket, get_blob, md5_hash, upload_from_*) is unchanged in 3.x. Follow-on fixes: - workers: bound `requires-python` to <3.13 like every other project. The open bound made uv resolve Python 3.14, where google-api-core >= 2.27 needs protobuf >= 6.31 but the connectors' secret-manager / bigquery pins cap it below 5. The image runs Python 3.12. - MinioFS: drop the UN-3487 `walk()` override. fsspec 2026.x's `DirFileSystem.walk()` relpaths each entry's `name` itself, so the override relpathed twice and tripped fsspec's assertion. The UN-3487 regression test now guards the upstream behaviour. - cohere embed timeout patch: re-pointed to 1.104.0 after re-diffing upstream; the sync `embedding()` still builds an untimed HTTPHandler. Known gap, deliberately not fixed here: boto3 >= 1.36 stops sending Content-MD5 on DeleteObjects, which MinIO older than RELEASE.2025-01-20 rejects. Verified live against the enterprise chart's bitnami-minio:2024.12.18: single-file deletes (recursive=False, raw fsspec) fail with MissingContentMD5, directory cleanup survives through the UN-3421 fallback. MinIO 2026-09-22 (OSS compose) is unaffected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…than 2025-01-20 botocore 1.36+ sends a CRC32 flexible checksum instead of Content-MD5 on DeleteObjects. MinIO older than RELEASE.2025-01-20 -- including the enterprise chart's bitnami-minio:2024.12.18 -- rejects that with MissingContentMD5, and s3fs routes every `rm` through DeleteObjects, so with the boto3 1.43 upgrade single-file deletes (recursive=False, raw fsspec) and connector deletes failed on those servers. The UN-3421 fallback only covered recursive deletes through sdk1 FileStorage. Register a `before-call.s3.DeleteObjects` handler that adds Content-MD5 over the final (XML-escaped) body. It is appended to botocore's BUILTIN_HANDLERS, so every session created afterwards gets it -- boto3, botocore and the aiobotocore sessions s3fs creates. AWS S3 and new MinIO accept both headers; S3 Express (which rejects MD5) is skipped. `AWS_REQUEST_CHECKSUM_CALCULATION=when_required` does not help (the operation requires a checksum), and botocore's deprecated `conditionally_calculate_md5` skips requests that carry a flexible checksum, so neither could be reused. unstract-connectors does not depend on sdk1, so each carries an identical, idempotent copy: sdk1 loads it from the file-storage helper, connectors from the MinIO connector. Whichever loads first registers. Verified live on the new stack against bitnami-minio 2024.12.18 and MinIO 2026-09-22: FileStorage.rm(recursive=False), raw fsspec single and bulk rm, and MinioFS rm (file and recursive) all succeed with no fallback warnings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
SonarCloud failed the quality gate on duplicated new code (39.2% vs 3%): connectors carried a verbatim copy of sdk1's hook and its tests, on the assumption that connectors does not depend on sdk1. It does at runtime: `unstract_file_system.py`, the base class every filesystem connector extends, imports `unstract.filesystem`, which depends on sdk1. So the MinIO connector now imports `unstract.sdk1.patches.s3_delete_objects_md5` directly, and the copy and its duplicate tests are gone. Importing the connector still registers the handler (verified), and connector deletes against MinIO 2024-12-18 still succeed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Deepak-Kesavan
left a comment
There was a problem hiding this comment.
🤖 Automated review
Automated review by Unstract PR review kit (Claude Code). Each finding below was reproduced against the code rather than inferred, so treat it as something to resolve before merge. If one is wrong, disagree on the thread and close it — that is the expected way to clear a finding. Anything tagged [unverified] was not reproduced and is flagged for your judgement instead.
4 inline comment(s) · 1 finding(s) not on changed lines (below).
Findings off the diff
unstract/sdk1/src/unstract/sdk1/adapters/base1.py:396
[major] The GPT-6 sampling strip covers the OpenAI, Azure OpenAI and Bedrock adapters but not the OpenAI-Compatible adapter. Its reasoning-model detector (^(o1|o3|o4|gpt-5)) was not extended togpt-6, so GPT-6 calls through it still sendtemperature=0.1andmax_tokens.
A user sets up the OpenAI-Compatible adapter (custom_openai) with api_base=https://api.openai.com/v1 or a LiteLLM/OpenAI gateway, and model=gpt-6-luna. OpenAICompatibleLLMParameters.validate (base1.py:438-529) never calls _strip_deprecated_sampling_params. Its only temperature drop is the reasoning branch at base1.py:501-517, and _is_openai_reasoning_model returns False for gpt-6. So every complete/acomplete/stream call sends temperature=0.1 plus a top-level max_tokens. That is the same request the PR says GPT-6 rejects, so every call to that adapter fails with a 400. Per-call temperature from the cloud agentic workers (LLM.complete(temperature=...)) also goes straight through, because _revalidate only re-runs this same validate.
Suggested fix: Add gpt-6 to _OPENAI_REASONING_MODEL_PATTERN (^(o1|o3|o4|gpt-5|gpt-6)(?:[-/.]|$)) so auto-detected GPT-6 goes through the existing extra_body path, which also drops temperature and max_tokens. Also return _strip_deprecated_sampling_params(validated) from OpenAICompatibleLLMParameters.validate, as the OpenAI adapter now does. Add a test that validates custom_openai + gpt-6-luna and asserts no top-level temperature.
🤖 Unstract PR review kit (Claude Code) · review-pr-bot:8f836231df38
review-pr-bot:review
…ver GPT-6 gateways Review findings from the PR review kit (Deepak), all reproduced first. Storage (sdk1 `patches/storage_compat.py`, renamed from `s3_delete_objects_md5.py`): - [major] botocore >= 1.36 defaults `request_checksum_calculation` to `when_supported`, so over HTTPS every PutObject / UploadPart went out as `Content-Encoding: aws-chunked` with a CRC32 trailer -- including to Unstract Cloud Storage (GCS's S3 interop API) and TLS MinIO. Captured on the wire for boto3 put_object / upload_fileobj and s3fs pipe_file. `S3_CHECKSUM_CONFIG` (`when_required` for requests and responses) restores boto3 1.34's plain payload-signed requests and is now passed explicitly to every S3 client we create: sdk1 FileStorage (explicit config_kwargs still win), MinioFS (and so UCS), and the two boto3 clients in `backend/connector_v2/unstract_account.py`. DeleteObjects still requires a checksum, so the Content-MD5 hook stays. - [minor] gcsfs 2026.x swaps `GCSFileSystem` (and fsspec's `gcs`) for the experimental gRPC-backed `ExtendedGcsFileSystem` unless GCSFS_EXPERIMENTAL_ZB_HNS_SUPPORT=false. Defaulted (operators can still opt in) before gcsfs is imported; the GCS connector loads the module too. LLM: - [major] The OpenAI-compatible adapter never stripped GPT-6 sampling params. `gpt-6` joins the reasoning-model detector (which also now accepts `.` so `gpt-6.1-sol` and `gpt-5.6-*` match), and validate() returns through `_strip_deprecated_sampling_params` for unrecognised aliases. Tests: - [minor] Azure GPT-6 is now tested through a real LLM on all four paths (complete, complete_vision, stream_complete, acomplete) with LiteLLM captured, replacing the helper that hand-copied LLM.__init__. Both mutations from the review (call sites bypassing `_revalidate`; per-call kwargs applied after `cost_model`) now fail. - [minor] Fresh-interpreter wiring tests: importing only the sdk1 storage helper, or only the MinIO / GCS connector, applies the settings. Wire tests through FileStorage -> s3fs -> aiobotocore on an HTTPS endpoint pin plain uploads and Content-MD5 on DeleteObjects. Verified live against MinIO 2024-12-18 (chart) and 2026-09-22 (OSS): sdk1 and connector writes, reads and every delete path succeed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CI's unit-sdk1 group syncs only the `test` dependency group (no extras), so `tests/test_storage_compat.py` failed to import aiobotocore. Its wire tests drive FileStorage -> s3fs -> aiobotocore, and its wiring tests check gcsfs's class selection; at runtime both come from the `aws` / `gcs` extras. Add them to the test group at the extras' versions, as was done for boto3. Reproduced CI's environment locally (`uv sync --group test`, no extras): 760 passed, 2 deselected (slow/integration). Lockfiles relocked; the diff is the dependency-group metadata only, no package version changes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Unstract test resultsPer-group results
Critical paths
|
* UN-4224 [FIX] Fix GCS connector test connection on gcsfs 2026.x Since #2310 moved gcsfs from 2024.10 to 2026.10, test connection fails for every Google Cloud Storage connector with `HttpError: Required parameter: project, 400`. This is the rc.425 regression behind the staging "GCS ETL Pipeline" UI test (Submit stays disabled because test connection fails). `test_credentials()` called `info("/")`. gcsfs 2024.10 turned that into `GET b?project=<project>` (list buckets); gcsfs 2026.10 sends `GET b/` with no project, even when one is configured, and GCS rejects it. Use `ls("")`, which on 2026.10 sends exactly the old `GET b?project=...` request -- same permission, same behaviour. It is the only root-level `info` call on gcsfs in either repo. Verified against real GCS (project unstract-staging) with gcsfs 2026.10: the connector's test_credentials() now returns True; the rc.425 code fails with the same 400 seen in staging. A regression test records the gcsfs request and asserts a bucket listing that carries the project; it fails if `info("/")` comes back. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * UN-4224 [FIX] Make the GCS connector test independent of import order Review (Greptile): the test imported `gcsfs.GCSFileSystem`, which resolves to the experimental ExtendedGcsFileSystem when gcsfs is imported before storage_compat sets GCSFS_EXPERIMENTAL_ZB_HNS_SUPPORT=false -- e.g. when this file runs alone -- so it could exercise a different client from production. Import `gcsfs.core.GCSFileSystem`, the class production runs, and pin it with a test. Passes on its own with the env var unset and even with the experimental class forced on. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>



What
temperatureto OpenAI GPT-6 models (gpt-6-luna,gpt-6-sol,gpt-6-astra,gpt-6.1-sol) from the AWS Bedrock, native OpenAI and Azure OpenAI LLM adapters. Azure AI Foundry is covered through the existing strip.custom_openai) adapter now also detects GPT-6, and dotted names likegpt-6.1-solandgpt-5.6-*, as reasoning models, and strips sampling params for unrecognised aliases (review finding).Why
GPT-6 Luna on AWS Bedrock failed every call with
temperature not permitted for this model. GPT-6 models reject the parameter.Our adapters always pass a temperature: the pydantic default (0.1, or 1 for Azure OpenAI), and 1 when reasoning / extended thinking is on.
What reaches AWS depends on the route. I captured the outgoing request on LiteLLM 1.96.2 with
drop_params=True:openai.gpt-6-luna"temperature": 0.1, the reported failureus./global.openai.gpt-6-lunaopenai.gpt-5.6-terraGPT-5.6 worked only because LiteLLM drops a non-1 temperature for
gpt-5names. That rule doesn't match GPT-6.Not a regression from UN-4020 [FIX] Route AWS Bedrock Mantle models (GPT-5.6 Terra) via bedrock_mantle #2248 (UN-4020). That PR didn't handle temperature for any Mantle model.
How
_SAMPLING_DEPRECATED_MODEL_STEMSgets agpt-6stem. This strip already handles Claude Opus 4.7+ and is called by the Bedrock, Azure AI Foundry, Vertex and Anthropic adapters..→-normalisation matches every Bedrock encoding:bedrock/,bedrock_mantle/,us./global.profiles andgpt-6.1-*.gpt-5.6-*(which normalises togpt-5-6-*),gpt-60andgpt-oss-*.cost_model.LLMsetscost_modelaside at construction. A newLLM._revalidatehelper, used by all four completion paths, passes it back into re-validation. It goes in after the per-call kwargs, so a caller passingtemperature=cannot displace it. The cloud agentic workers do passtemperature=; this was flagged by review.Can this PR break any existing features. If yes, please list possible items. If no, please explain why. (PS: Admins do not merge the PR without this section filled)
gpt-60and older Claude models to keep their temperature.LLM._revalidatepassescost_modelback toadapter.validate(). No parameter model declares it, Pydantic ignores unknown keys, and every call site already popscost_modelbefore LiteLLM, so no other adapter's request changes.Database Migrations
Env Config
Relevant Docs
Related Issues or PRs
Dependencies Versions
LiteLLM 1.104.0 is the first release whose bundled registry includes the GPT-6 Bedrock Mantle entries. With it, air-gapped deployments now route
openai.gpt-6-lunato Mantle and price GPT-6 calls instead of recording $0. LiteLLM ≥ 1.100 hard-requiresboto3>=1.43.1, so the storage stack moves with it:Follow-on changes:
requires-pythonis now bounded to<3.13, like every other project. The open bound made uv resolve Python 3.14, where protobuf ≥ 6.31 conflicts with the connectors' secret-manager/bigqueryprotobuf<5caps. The image runs 3.12.walk()override is removed. fsspec 2026.x'sDirFileSystem.walk()now relpaths entry names itself, so the override relpathed twice and tripped fsspec's assertion.azure-datalake-store(ADLS Gen1) drops out with adlfs 2026; we only useAzureBlobFileSystem.Upload wire format (review finding). botocore ≥ 1.36 defaults
request_checksum_calculationtowhen_supported. Over HTTPS, every PutObject / UploadPart then goes out asaws-chunkedwith a CRC32 trailer. That includes Unstract Cloud Storage (GCS's S3 interop API) and TLS MinIO.storage_compat.S3_CHECKSUM_CONFIG(when_requiredfor requests and responses) restores boto3 1.34's plain payload-signed requests. It's passed explicitly to every S3 client we create: sdk1FileStorage(explicitconfig_kwargsstill win),MinioFSand therefore UCS, and the two boto3 clients inbackend/connector_v2/unstract_account.py. A real UCS upload against GCS has not been tested (no credentials); the request now matches what 1.34 sent.gcsfs experimental class (review finding). gcsfs 2026.x swaps
GCSFileSystem(and fsspec'sgcs) for the experimental gRPC-backedExtendedGcsFileSystemunlessGCSFS_EXPERIMENTAL_ZB_HNS_SUPPORT=false.storage_compatdefaults that tofalsebefore gcsfs is imported; an explicit operator setting still wins.MinIO older than RELEASE.2025-01-20. boto3 ≥ 1.36 sends a CRC32 checksum instead of
Content-MD5onDeleteObjects, and older MinIO rejects that withMissingContentMD5. That includes the enterprise chart'sbitnami-minio:2024.12.18.s3fsroutes everyrmthroughDeleteObjects.before-call.s3.DeleteObjectshandler re-addsContent-MD5over the final (XML-escaped) body. It is appended to botocore'sBUILTIN_HANDLERS, so boto3, botocore and the aiobotocore sessions thats3fscreates all get it. S3 Express, which rejects MD5, is skipped.unstract/sdk1/patches/storage_compat.py, loaded by the sdk1 file-storage helper and by the MinIO and GCS connectors. Connectors already depends on sdk1 at runtime, viaunstract.filesystem.AWS_REQUEST_CHECKSUM_CALCULATION=when_requireddoesn't help, becauseDeleteObjectsrequires a checksum. botocore's deprecatedconditionally_calculate_md5skips requests that carry a flexible checksum.Notes on Testing
bitnami-minio:2024.12.18and MinIO 2026-09-22, with the hook:FileStorage.rm(recursive=False), raw fsspec single and bulkrm, andMinioFSrm (file and recursive) all succeed with no fallback warnings. Without the hook, the single-file and raw-fsspec deletes fail on 2024.12.18.MinioFSls/walk names stay bucket-relative on both.DeleteObjectsrequest: the digest matches the XML-escaped body, the handler runs afterescape_xml_payload, and registration is idempotent. With the registration removed, 3 tests fail.bitnami-minio:2024.12.18:MissingContentMD5.Bulk delete failed with MissingContentMD5in the logs; it only completed because of the UN-3421 fallback. New tests intests/test_sampling_strip.pycover:LLM._revalidate, including a per-calltemperature=0.5, a deployment that names the model, and non-GPT-6 controls (which keep a per-call temperature);openai.gpt-6-lunaon Mantle: 200, completion returned (25/32 tokens). This is the route that failed before.global.openai.gpt-6-lunaandus.openai.gpt-6-lunaon Converse: 200.openai.gpt-6-lunaon Mantle returns "model does not exist" inap-south-1. That's regional availability, not this change.ruff0.3.4 check and format are clean.Screenshots
Checklist
I have read and understood the Contribution Guidelines.
🤖 Generated with Claude Code