Aggregating data frames
Data frames can be aggregated as a whole or over groups of rows. The modern dataframe workflow uses immutable aggregation descriptors created by the static Aggregation class. Older AggregatorGroup and AggregateBy APIs remain available and are still useful when working directly with the lower-level aggregation primitives.
Aggregate with descriptors
An aggregation descriptor describes a request: the selected column or columns, the aggregator to apply, and the output column key. Descriptors do not execute by themselves. They are bound and executed when passed to a dataframe Aggregate overload.
Default unary descriptor output keys are strings in column_statistic form:
var summary = frame.Aggregate(
Aggregation.Sum("sales"),
Aggregation.Mean("margin"));Aggregate selected columns
Descriptor helpers can aggregate all columns selected by a ColumnSelector. Selector-based helpers require the source column-key type because ColumnSelector itself is non-generic.
var numericSummary = frame.Aggregate(
Aggregation.Mean<string>(ColumnSelector.OfType<double>()),
Aggregation.StandardDeviation<string>(ColumnSelector.OfType<double>()));Role metadata can be used to define reusable semantic selections for aggregation. Roles are stored on the data frame and reflected in schema snapshots.
frame.SetColumnRole(ColumnSelector.OfType<double>(), ColumnRole.Feature);
frame.SetColumnRole("target", ColumnRole.Target);
var featureSummary = frame.Aggregate(
Aggregation.Mean<string>(ColumnSelector.HasRole(ColumnRole.Feature)));Aggregate pairs of columns
Binary descriptor helpers aggregate pairs of selected columns. Default binary output keys are strings in left_right_statistic form. When binary selectors resolve to multiple columns, they expand cartesian: every left selected column is paired with every right selected column.
var relationship = frame.Aggregate(
Aggregation.Correlation("height", "weight"));Aggregate over groups
Descriptor aggregation also works with existing Grouping<TKey> objects and dataframe grouping templates. Grouped descriptor aggregation uses the grouping index as the result row index.
var grouped = frame.Aggregate(
Grouping.By<string, string>("region"),
Aggregation.Sum("sales"),
Aggregation.Mean("margin"));Choose output keys
Use As<COut2> to choose the output key produced by a descriptor. This is the path for custom string names and for typed result column keys. All descriptors in one Aggregate call must share the same output-key type. Duplicate output keys throw.
var namedSummary = frame.Aggregate(
Aggregation.Sum("sales").As("totalSales"));Use aggregator groups directly
The existing AggregatorGroup and AggregateBy APIs remain available. They expose the lower-level aggregation primitives directly and are still accurate for code that aggregates every applicable column with the same aggregator group.
For dataframe reporting workflows, descriptor aggregation is usually clearer because each output column is represented by an explicit request and a predictable output key. Descriptor aggregation can also combine unary and binary outputs naturally because the default output key type is String.
The following older examples use aggregator groups directly:
var titanic = DelimitedTextFile.ReadDataFrame(titanicFilename);
var means = titanic.Aggregate(Aggregators.Mean.As<double>());