我有一个使用XML文件构建的数据框,现在我想统计并求和其值,类似于SQL中的计数和求和。
这就是数据框的样子:
msgDataSource msgFileSource processDate msgNumRows
1 source1 Quarter 2015-01-30 30
2 source1 Month 2015-01-30 15
3 source1 Month 2015-01-30 20
4 source1 Year 2015-01-30 1
5 source2 Quarter 2015-01-30 30
6 source3 Quarter 2015-01-30 15
7 source1 Year 2015-02-01 80
8 source2 Year 2015-02-01 90
9 source1 Quarter 2015-02-01 5
10 source2 Quarter 2015-03-15 9
11 source3 Quarter 2015-03-15 14
这是我需要的内容
processDate msgFileSource msgDataSource sumDataSources countDataSources
1: 2015-01-30 Month source1 35 2
2: 2015-01-30 Quarter source1 30 1
3: 2015-01-30 Quarter source2 30 1
4: 2015-01-30 Quarter source3 15 1
5: 2015-01-30 Year source1 1 1
6: 2015-02-01 Quarter source1 5 1
7: 2015-02-01 Year source1 80 1
8: 2015-02-01 Year source2 90 1
9: 2015-03-15 Quarter source2 9 1
10: 2015-03-15 Quarter source3 14 1
这是我目前能够得到的内容:
目前为止,这就是我能够得到的。
processDate msgFileSource msgDataSource sumDataSources
1: 2015-01-30 Month source1 35
2: 2015-01-30 Quarter source1 30
3: 2015-01-30 Quarter source2 30
4: 2015-01-30 Quarter source3 15
5: 2015-01-30 Year source1 1
6: 2015-02-01 Quarter source1 5
7: 2015-02-01 Year source1 80
8: 2015-02-01 Year source2 90
9: 2015-03-15 Quarter source2 9
10: 2015-03-15 Quarter source3 14
这是我的代码:
dfFullData <- data.frame (
msgDataSource = c("source1", "source1", "source1", "source1", "source2", "source3", "source1", "source2", "source1", "source2", "source3"),
msgFileSource = c("Quarter", "Month", "Month", "Year", "Quarter", "Quarter", "Year", "Year", "Quarter", "Quarter", "Quarter"),
processDate = c("2015-01-30", "2015-01-30", "2015-01-30", "2015-01-30", "2015-01-30", "2015-01-30", "2015-02-01", "2015-02-01", "2015-02-01", "2015-03-15", "2015-03-15"),
msgNumRows = c(30, 15, 20, 1, 30, 15, 80, 90, 5, 9, 14),
stringsAsFactors=FALSE
)
summaryTable <- data.table(dfFullData)
summaryTable <- summaryTable[
order(processDate, msgFileSource, msgDataSource),
sum(msgNumRows),
by=list(processDate, msgFileSource, msgDataSource)
]
setnames(summaryTable, "V1", "sumDataSources")
print(summaryTable)
有没有一种方法可以在一次操作中计算数量,或者我应该分别计算然后执行cbind?
我怎样才能达到我需要的效果呢?
谢谢。
order()
有什么原因吗?另外,length(.)
只是.N
-特殊的内置符号。 - Arunkeyby
代替by
,而不是使用order()
-keyby
将按分组列对数据进行排序后聚合 - 这更有效,因为它在聚合数据上排序。有关更多信息,请查看这些新的HTML小品。 - Arun