Skip to content
View STEFANOVIVAS's full-sized avatar

Block or report STEFANOVIVAS

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
stefanovivas/README.md

👋 Hi there! Welcome to my GitHub profile.

  • I’m currently working at Caixa Econômica Federal as a mid-level Data Engineer.
  • The tools I use in my day-to-day work are Databricks,SQL Server, SSIS, Power BI, and Python.
  • I have a bachelor's degree in Economics from the Federal University of Pernambuco(UFPE) and an MBA in data engineering.
  • I know Python, SQL, and Java languages.
  • I’m interested in all the Data stack, Financial Markets, and Homebrewing.
  • How to reach me: www.linkedin.com/in/stéfano-vivas

🎉 Open Source Contributions

Databrickslabs DQX (data quality framework)

New features proposed:

  • Add ability to profile and generate rules on subset of the input data by introducing a sql expression filter (#569)
  • Improve summary stats report for string datatype columns (#670)
  • Extend summary metrics with per-check-name breakdowns (#943)
  • Introduces a set of checks called is_null, is_empty, and is_null_or_empty to allow checks that confirm whether a given column is null, empty, or both (#965)
  • Added profile build for "has_no_outliers" check (#977)
  • Extended dq_generate_min_max method to support Python's Decimal type in addition to int and float for min/max validation checks (#1013)

New features implemented:

  • Introduces a new data quality check to detect outliers for numeric columns using the Median Absolute Deviation method (MAD) (#359)
  • Add ability to profile and generate rules on subset of the input data by introducing a sql expression filter (#569)
  • Added rule_fingerprint, rule_set_fingerprint, and created_at to checks storage, so now it can be versioned (#672)
  • Introduces a new dataset-level quality check, aggr_matches_dataset, to validate ingestion correctness by comparing an aggregate metric (row count by default, or other curated/built-in aggregates) computed on the checked DataFrame against the same metric computed on an dataset reference (#1309)

Pinned Loading

  1. azure-databricks-project azure-databricks-project Public

    Formula one data engineer project with Azure, databricks and Spark

    Python

  2. rais-pipeline rais-pipeline Public

    A pipeline for extract, load and transform Brazilian job market dataset RAIS

    Python

  3. Investiment-funds-de-project Investiment-funds-de-project Public

    Data engineer project to create a dashboard with investment funds quotas from Microsoft azure tools.