You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
to_sql() ignores the connection's botocore config and the credentials of connect(session=...) #1067
pyathena.pandas.util.to_sql() builds its own S3 clients instead of using the connection's settings:
botocore config. The bucket resource (for if_exists="replace") and the upload workers' resources (to_parquet()) are created without config, so neither the connection's config nor its s3_config (Share one S3 client per connection across result sets and filesystems #1058) applies: proxies, timeouts, retries, max_pool_connections and PyAthena's user agent are missing from these requests.
Credentials of connect(session=...). The upload workers create Session(**conn._session_kwargs, profile_name=conn.profile_name). A session passed to connect() is not among these arguments, so the uploads use the default credential chain, while the bucket resource uses the connection's session. One to_sql() call can then delete with one identity and upload with another, or fail to upload.
Checked locally on master 911492c with a session that has explicit keys:
fromcopyimportdeepcopyfromboto3.sessionimportSessionfrompyathenaimportconnectuser_session=Session(aws_access_key_id="USER_KEY", aws_secret_access_key="s", region_name="us-west-2")
conn=connect(s3_staging_dir="s3://bucket/path/", region_name="us-west-2", session=user_session)
session_kwargs=deepcopy(conn._session_kwargs)
session_kwargs.update({"profile_name": conn.profile_name})
print(session_kwargs) # {'profile_name': None}print(Session(**session_kwargs).get_credentials().access_key) # the default chain's key, not USER_KEY
role_arn and serial_number are not affected: Connection stores the resulting temporary credentials in _kwargs, which reach the workers.
Expected
The S3 requests of to_sql() use the connection's s3_config and the credentials of the connection's session, with both ThreadPoolExecutor and ProcessPoolExecutor (the worker arguments must stay picklable; botocore.config.Config pickles).
Follow-up of #1058, which made to_sql() leave out Athena's endpoint_url and api_version.
Problem
pyathena.pandas.util.to_sql()builds its own S3 clients instead of using the connection's settings:if_exists="replace") and the upload workers' resources (to_parquet()) are created withoutconfig, so neither the connection'sconfignor itss3_config(Share one S3 client per connection across result sets and filesystems #1058) applies: proxies, timeouts, retries,max_pool_connectionsand PyAthena's user agent are missing from these requests.connect(session=...). The upload workers createSession(**conn._session_kwargs, profile_name=conn.profile_name). Asessionpassed toconnect()is not among these arguments, so the uploads use the default credential chain, while the bucket resource uses the connection's session. Oneto_sql()call can then delete with one identity and upload with another, or fail to upload.Checked locally on master 911492c with a session that has explicit keys:
role_arnandserial_numberare not affected:Connectionstores the resulting temporary credentials in_kwargs, which reach the workers.Expected
The S3 requests of
to_sql()use the connection'ss3_configand the credentials of the connection's session, with bothThreadPoolExecutorandProcessPoolExecutor(the worker arguments must stay picklable;botocore.config.Configpickles).Follow-up of #1058, which made
to_sql()leave out Athena'sendpoint_urlandapi_version.