Skip to content

.bruin.yml Reference ​

The .bruin.yml file is the central configuration file for Bruin pipelines. It stores all credentials, connection details, and environment configurations needed to run your data pipelines. The file is automatically created when you run any bruin command the first time in a project, and it is automatically added to .gitignore.

.bruin.yml file is expected to be in the root of your Git repo. You can specify a different location using the --config-file /path/to/.bruin.yml

File Structure ​

The file is a YAML file with the following structure:

yaml
default_environment: <environment_name>
environments:
  <environment_name>:
    schema_prefix: <optional_prefix>
    config:
      full_refresh_restricted: true
    connections:
      <connection_type>:
        - name: "<connection_name>"
          # connection-specific fields...

Example ​

yaml
default_environment: default
environments:
  default:
    connections:
      duckdb:
        - name: "duckdb-default"
          path: "duckdb.db"
      chess:
        - name: "chess-default"
          players:
            - "FabianoCaruana"
            - "Hikaru"
            - "MagnusCarlsen"
            - "GothamChess"
            - "DanielNaroditsky"
            - "AnishGiri"
            - "Firouzja2003"
            - "LevonAronian"
            - "WesleySo"
            - "GarryKasparov"

Top-Level Fields ​

FieldTypeRequiredDescription
default_environmentstringNoEnvironment to use when none is specified. Defaults to default.
environmentsmapYesMap of environment names to their configurations.

Environment Configuration ​

Each environment contains:

FieldTypeRequiredDescription
connectionsobjectYesConnection definitions grouped by type.
schema_prefixstringNoPrefix added to schema names (useful for dev/staging environments).
config.full_refresh_restrictedbooleanNoPrevents --full-refresh from dropping and recreating tables for all assets in this environment.

Environment Variables ​

You can reference environment variables in your configuration using ${VAR_NAME} syntax:

yaml
environments:
  default:
    connections:
      postgres:
        - name: my_postgres
          username: ${POSTGRES_USERNAME}
          password: ${POSTGRES_PASSWORD}
          host: ${POSTGRES_HOST}
          port: ${POSTGRES_PORT}
          database: ${POSTGRES_DATABASE}

Environment variables are expanded at runtime, not when the file is parsed.

Generic Credentials ​

Generic credentials are key-value pairs that can be used to inject secrets into your assets:

yaml
connections:
  generic:
    - name: MY_SECRET
      value: secretvalue

Common use cases include API keys, passwords, and other secrets that don't fit a specific connection type.

Connection Concurrency Limits ​

Connections can define max_concurrent_assets to cap how many assets can use that connection at the same time in a single pipeline run:

yaml
environments:
  default:
    connections:
      postgres:
        - name: "postgres-main"
          username: "${POSTGRES_USER}"
          password: "${POSTGRES_PASSWORD}"
          host: "db.example.com"
          port: 5432
          database: "analytics"
          max_concurrent_assets: 2

The value must be a positive integer. When the limit is reached, Bruin waits to start additional assets that need the same connection until another asset using that connection finishes.

This limit is separate from the --workers setting. --workers controls total asset concurrency for a run, while max_concurrent_assets controls concurrency for one named connection. For more detail, see Concurrency & Resource Limits.

Connection Types ​

For the specific fields and configuration options for each connection type, refer to the dedicated documentation pages:

Data Platforms ​

Connection TypeDocumentation
google_cloud_platformGoogle BigQuery
snowflakeSnowflake
postgresPostgreSQL
redshiftRedshift
databricksDatabricks
athenaAWS Athena
duckdbDuckDB
motherduckMotherDuck
clickhouseClickHouse
mysqlMySQL
dorisApache Doris
mssqlMicrosoft SQL Server
synapseAzure Synapse
oracleOracle
trinoTrino
elasticsearchElasticsearch
mongo_atlasMongoDB Atlas
s3S3
gcsGoogle Cloud Storage
emr_serverlessAWS EMR Serverless
dataproc_serverlessGCP Dataproc Serverless

Data Ingestion Sources ​

Connection schemas for ingestion sources are documented on their respective pages under Data Ingestion. Each source page includes the required .bruin.yml connection configuration.


Complete Example ​

yaml
default_environment: default

environments:
  default:
    connections:
      google_cloud_platform:
        - name: "gcp-prod"
          project_id: "my-project"
          service_account_file: "credentials/gcp-service-account.json"

      postgres:
        - name: "postgres-main"
          username: "${POSTGRES_USER}"
          password: "${POSTGRES_PASSWORD}"
          host: "db.example.com"
          port: 5432
          database: "analytics"

      snowflake:
        - name: "snowflake-prod"
          account: "ABC12345"
          username: "bruin_user"
          private_key_path: "credentials/snowflake_key.p8"
          database: "ANALYTICS"
          warehouse: "COMPUTE_WH"

      generic:
        - name: "SLACK_WEBHOOK"
          value: "https://hooks.slack.com/..."

  staging:
    schema_prefix: "stg_"
    connections:
      google_cloud_platform:
        - name: "gcp-staging"
          project_id: "my-project-staging"
          use_application_default_credentials: true

      postgres:
        - name: "postgres-staging"
          username: "staging_user"
          password: "${STAGING_POSTGRES_PASSWORD}"
          host: "staging-db.example.com"
          port: 5432
          database: "analytics"