Empty job artifact fails with 500: Workhorse sends CompleteMultipartUpload without any parts

When a job uploads an artifact file of exactly 0 bytes and artifacts go to object storage via direct upload, Workhorse breaks out of its multipart loop before it appends a single part, but still issues CompleteMultipartUpload. The request goes out with an empty part list, S3 rejects it with 400, Workhorse answers the Runner with 500, and the Runner retries.

Steps to reproduce

Artifacts on S3-compatible object storage (artifacts_object_store_enabled = true, path_style = true, aws_signature_version = 4), then a job that produces an empty artifact file:

codequality:
  script:
    - touch gl-code-quality-report.json
  artifacts:
    reports:
      codequality: gl-code-quality-report.json

Any artifact type reaches the same code path, reports:codequality is just where we hit it, with a job that produced an empty report file.

What happens

Workhorse log, endpoint and object path redacted:

{"error":"handleFileUploads: extract files from multipart: persisting multipart file: CompleteMultipartUpload request https://<s3-endpoint>/artifacts/<hash>/%40final/<path>?uploadId=<id>&X-Amz-Expires=15300&X-Amz-Date=<ts>&X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-SignedHeaders=content-type%3Bhost&X-Amz-Signature=[FILTERED] returned: 400 Bad Request","level":"error","method":"POST","time":"<ts>","uri":"/api/v4/jobs/<id>/artifacts?artifact_format=raw&artifact_type=codequality&expire_in=1+year"}

Access log line for the same request:

{"method":"POST","route":"^/api/v4/jobs/[0-9]+/artifacts\\z","status":500,"read_bytes":829,"written_bytes":22,"duration_ms":295}

read_bytes is the entire multipart body. Across our three Workhorse nodes in the same period, 174 successful uploads of that artifact type had read_bytes of 1018 or more, and everything down at the 808 to 831 byte mark failed, which is the multipart form overhead with nothing in the file part.

Why

workhorse/internal/upload/destination/objectstore/multipart.go, v19.3.1 and unchanged on master:

func (m *Multipart) readAndUploadOnePart(...) (*s3api.CompleteMultipartUploadPart, error) {
	...
	n, err := io.Copy(file, src)
	...
	if n == 0 {
		return nil, nil
	}
	for i, partURL := range m.PartURLs {
		src := io.LimitReader(r, m.partSize)
		part, err := m.readAndUploadOnePart(ctx, partURL, m.PutHeaders, src, i+1)
		if err != nil {
			return err
		}
		if part == nil {
			break
		}
		cmu.Part = append(cmu.Part, part)
	}
	...
	if err := m.complete(ctx, cmu); err != nil {
		return err
	}

With an empty file the first part reads nothing, readAndUploadOnePart returns nil, nil, the loop breaks before the append, and cmu.Part stays empty. Nothing guards complete() against that and there is no PutObject fallback, so the XML goes out without a <Part> element. S3 wants at least one.

A single byte in the file is enough to avoid it. Then n > 0, one part gets appended, and the upload succeeds, since there is no minimum size for the only part.

Expected

Either store the empty object, or fail with a 4xx that names the cause.

The URL for the first option is already in the response: ObjectStorage::DirectUpload#to_hash emits StoreURL alongside MultipartUpload. Workhorse cannot reach it at the point of failure though, because the Multipart strategy only carries PartURLs, CompleteURL, AbortURL and DeleteURL. A fallback would have to sit in objectstore/uploader.go where the strategy is picked, not in multipart.go, which makes this a question of intended behaviour rather than a local guard clause.

The 500 is what makes the current behaviour expensive: the Runner treats it as retryable and repeats the upload, so one bad artifact turns into a burst of failures spread across every Workhorse in the cluster. Three such jobs took the WorkhorseHighErrorRate alert for ^/api/v4/jobs/[0-9]+/artifacts\z to 78.9% on one of our nodes, and the actual cause is only visible in the Workhorse log of whichever node happened to serve the request.

I can send an MR with a test case for multipart_test.go once you have settled which of the two you want.

Environment

GitLab 19.3.1-ee, Omnibus, gitlab-workhorse v19.3.1-ee, Runner 19.2.1, S3-compatible object storage.

Edited by roth-wine