FindAlternative
Back to Home
fg-data-profiling

fg-data-profiling

One-line data quality profiling for Pandas and Spark DataFrames

softwareAnalytics & BIdata profilingpandasspark
Our Verdict

Best for

Python data scientists needing quick, code‑only profiling of Pandas or Spark data

Skip if

Those requiring a GUI, cross‑language support, or built‑in scheduling

What is fg-data-profiling?

fg-data-profiling provides a minimal‑code way to generate comprehensive data quality reports for both Pandas and Spark DataFrames. With a single function call you obtain statistics, missing‑value analysis, and visual summaries that help you understand your data quickly. The library is open‑source, lightweight, and integrates seamlessly into existing Python data pipelines, making it ideal for exploratory data analysis, data cleaning, and model preparation without heavy configuration.

SpecificationsAI-estimated

deploymentSelf-hosted
open source✅ Yes
github stars13,660
api available✅ Yes
support optionsGitHub Issues
key integrationsPandas, PySpark
primary languagePython

Key Features of fg-data-profiling

Generate a full profiling report with a single function call on a Pandas DataFrame.
Generate the same comprehensive report on a Spark DataFrame without code changes.
Automatically compute column‑wise statistics such as mean, median, min, max, and distinct counts.
Detect and summarize missing, null, and NaN values across all columns.
Provide histogram and distribution visualizations for numeric and categorical data.
Identify high‑cardinality columns and potential outliers in a concise table.
Export the profiling results to HTML or JSON for sharing with stakeholders.

Use Cases for fg-data-profiling

1

Quick EDA

Create an instant data quality overview before deeper analysis.

2

Data Cleaning

Spot missing and inconsistent values to prioritize cleaning steps.

3

Model Validation

Verify feature distributions and detect drift between training and new data.

4

Pipeline Monitoring

Integrate profiling into ETL jobs to catch data issues early.

Pros & Cons of fg-data-profiling

Pros

  • Zero‑configuration profiling with a single line of code.
  • Supports both Pandas and Spark DataFrames.
  • Lightweight and fast, suitable for large datasets.
  • Open‑source with active community contributions.

Cons

  • Limited to Python; no native UI beyond generated HTML.
  • Advanced visual customisation requires manual tweaking.
  • No built‑in scheduling; must be invoked from external pipelines.

Frequently Asked Questions

Does fg-data-profiling work with Spark on a cluster?

Yes, you can pass a Spark DataFrame and the library will compute the profile using Spark's distributed execution.

What output formats are available?

Reports can be saved as interactive HTML files or as JSON for programmatic consumption.

Is there a way to customize which statistics are calculated?

You can configure the profiling function with optional parameters to include or exclude specific metrics.

How is the library licensed?

fg-data-profiling is released under the Apache 2.0 license, allowing free commercial and private use.

Pricing Overview

View full pricing →
Free

Detailed plans are not listed. Visit the official website for pricing information.

No reviews yet. Be the first to write one!

Top Alternatives & Similar Tools

View all alternatives & similar tools →

No alternatives available yet.

People also viewed

Related searches

About the Tool

Unclaimed Listing
Socials
Target AudienceData scientists and analysts

Is this your tool?

Claim this page to update details, reply to user reviews, and drive more traffic to your product.

Claim this Product →

Tags

data profilingpandassparkEDAopen-source

Explore Related Topics

Build with AI

Discover AI tools to supercharge your workflow.

Explore AI tools