Skip to main content

Azure Blob Collector Job Discovers Zero Events When Source JSON Has a UTF-8 BOM

  • September 28, 2026
  • 0 replies
  • 3 views

Jessica Bracken

Symptom

An Azure Blob Storage Collector job discovers the target blob but extracts zero events from it. The following conditions are all present:

  • Discovery succeeds — the target blob appears in the Collector's discovery results.
  • Blob content reads succeed with no network, DNS, or authentication errors. clientSecret authentication with the Storage Blob Data Contributor role is sufficient for both list and read operations.
  • Swapping the job's pipeline to passthru makes no difference. The event count is still zero.
  • No error or warning appears anywhere, not in Job Inspector and not in the worker cribl.log at any log level. The task reports success with zero events.
  • The source file, when opened in a text editor or viewed in the Azure Portal, clearly contains valid JSON data.

Environment

  • Cribl Stream (Cloud-managed or self-managed). Not version-specific — applies to any release using the json_array Event Breaker type, or a built-in pipeline with equivalent array-detection logic (for example, azureblobjson).
  • Azure Blob Storage Collector source, or any Collector or source whose input breaker uses the json_array Event Breaker type.
  • Source file is UTF-8-encoded JSON with a leading byte order mark (BOM).

Resolution

Resolve the issue in two steps: confirm that a BOM is the cause, then remove the BOM at the source. Cribl Stream cannot detect or strip a BOM in the current product.

Confirm the Pattern

  1. Verify that discovery succeeds while the run reports zero discovered or extracted events.
  2. Verify that no error or warning appears in Job Inspector or in the worker cribl.log, at any log level, for the task.
  3. Set the job's pipeline to passthru.
  4. Run the job again. If the event count is still zero, the failure occurs upstream of the pipeline, in the input breaker's array-detection step.
  5. Inspect the source file for a leading UTF-8 BOM using one of the following methods:
    • Open the file in Notepad++ and check the Encoding menu. It displays UTF-8-BOM for an affected file, or plain UTF-8 for a clean one.
    • Run xxd <file> | head -1 from a command line. A BOM appears as ef bb bf immediately before 5b, the byte for [.
    • Run python3 -c "print(open('<file>','rb').read(3))". A BOM prints as b'\xef\xbb\xbf'.

Resolve the Issue at the Source

Identify and fix the process that produces the BOM-prefixed file. There is no in-product workaround for this defect.

  1. Identify the process that generates the source file. A PowerShell export script is the most common cause.
  2. Update the process to write the file without a UTF-8 BOM:
    • In PowerShell 7 and later, use -Encoding utf8NoBOM on Out-File or Set-Content.
    • In Windows PowerShell 5.1, Out-File and Set-Content have no BOM-less UTF-8 option. Use [System.IO.File]::WriteAllText($path, $jsonContent, [System.Text.UTF8Encoding]::new($false)) instead.
    • For a one-off manual correction, open the file in Notepad++, select the Encoding menu, select UTF-8 (not UTF-8-BOM), and save the file.
  3. Upload the corrected file to the same container or path.
  4. Run the Collector job again. Verify that the discovered event count now matches the number of real records in the file.

Cause

The json_array Event Breaker type, and equivalent logic in built-in pipelines such as azureblobjson, determines whether content is a JSON array using a naive prefix check. The check is effectively content.startsWith('[') rather than a full JSON parse.

A UTF-8 BOM (the 3 bytes EF BB BF) placed before the opening [ defeats this check. The check returns false, so the Event Breaker concludes the content is not a JSON array and discovers zero events, even though the file contains valid records.

The underlying data is valid JSON. Any standards-compliant JSON parser, including JSON.parse and Python's json.loads, strips or tolerates a leading BOM and parses the content correctly.

BOM-prefixed UTF-8 output is a common artifact of Windows PowerShell 5.1's Out-File -Encoding utf8 and Set-Content -Encoding utf8, both of which default to BOM-prefixed UTF-8. Export scripts against Azure, Microsoft Entra ID, and Microsoft Graph APIs commonly rely on this default. The pattern applies broadly to any Blob Storage, S3, or file-based Collector source fed by PowerShell-generated JSON.

No error or warning surfaces in the product for this condition. The Event Breaker treats "not a JSON array" as a normal outcome, not a failure, so the task reports success with zero events.

Additional Information

  • Related conditions ruled out while researching this issue:
    • DNS or private-endpoint reachability — no supporting log evidence; blob content reads succeeded consistently.
    • RBAC or permissions — the Storage Blob Data Contributor role already covers both list and blob-read data actions; no separate permission is required for blob content reads.
    • Unsupported AppendBlob blob type (a distinct pattern seen in other instances) — ruled out by confirming blobType: "BlockBlob" in the discovery task output.
    • Pipeline-level maxEventBytes mismatch — initially suspected due to real truncation warnings observed earlier in the same instance, but ruled out by the passthru pipeline test, since passthru has no Event Breaker function at all and the event count stayed at zero.
  • A related public community thread describes the same underlying pattern on a Filesystem Collector source (BOM-prefixed TSV file, event breaker mismatch resolved by removing the BOM in Notepad++): Filesystem Collector and Event Breaker Inconsistencies | Community. This confirms the pattern is not unique to Azure Blob Collector sources.