標籤:總結 extends 個數 text spark within serial isa log
rdd.join的實現:rdd1.join(rdd2) => rdd1.cogroup(rdd2,partitioner)
/** * Return an RDD containing all pairs of elements with matching keys in `this` and `other`. Each * pair of elements will be returned as a (k, (v1, v2)) tuple, where (k, v1) is in `this` and * (k, v2) is in `other`. Uses the given Partitioner to partition the output RDD. */ def join[W](other: RDD[(K, W)], partitioner: Partitioner): RDD[(K, (V, W))] = self.withScope {
//rdd.join的實現:rdd1.join(rdd2) => rdd1.cogroup(rdd2,partitioner) => flatMapValues(遍曆兩個value的迭代器)
//最後返回的是(key,(v1,v2))這種形式的元組
this.cogroup(other, partitioner).flatMapValues( pair => for (v <- pair._1.iterator; w <- pair._2.iterator) yield (v, w) ) }
跟到cogroup方法
/** * For each key k in `this` or `other`, return a resulting RDD that contains a tuple with the * list of values for that key in `this` as well as `other`. */ def cogroup[W](other: RDD[(K, W)], partitioner: Partitioner) : RDD[(K, (Iterable[V], Iterable[W]))] = self.withScope { if (partitioner.isInstanceOf[HashPartitioner] && keyClass.isArray) { throw new SparkException("Default partitioner cannot partition array keys.") } /** * 這裡構造一個CoGroupedRDD,也就是 cg = new CoGroupedRDD(Seq(rdd1,rdd2),partitioner) * 其索引值對中的value要求是Iterable[V]和Iterable[W]類型 * 下面瞭解CoGroupedRDD這個類,看是怎麼構造的 */ val cg = new CoGroupedRDD[K](Seq(self, other), partitioner) cg.mapValues { case Array(vs, w1s) => (vs.asInstanceOf[Iterable[V]], w1s.asInstanceOf[Iterable[W]]) } }
這是CoGroupedRDD的類聲明,其中有兩個與java 文法的不同:
1.型別宣告中的小於符號“<”,這個在scala 中叫做變數類型的上界,也就是原類型應該是右邊類型的子類型,具體參見《快學scala》的17.3節
[email protected]:這個是瞬時變數註解,不用進行序列化 ,也可以參見《快學Scala》的15.3節
/** 這裡返回的rdd的類型是(K,Array[Iterable[_]]),即key不變,value為所有對應這個key的value的迭代器的數組*/class CoGroupedRDD[K: ClassTag]( @transient var rdds: Seq[RDD[_ <: Product2[K, _]]], part: Partitioner) extends RDD[(K, Array[Iterable[_]])](rdds.head.context, Nil)
看看這個RDD的依賴以及如何分區的
再看這兩個函數之前,最好先瞭解下這兩個類是幹什麼的:
1.CoGroupPartition是Partition的一個子類,其narrowDeps是NarrowCoGroupSplitDep類型的一個數組
/** * 這裡說到CoGroupPartition 包含著父RDD依賴的映射關係, * @param index:可以看到CoGroupPartition 將index作為雜湊code進行分區 * @param narrowDeps:narrowDeps是窄依賴對應的分區數組 */private[spark] class CoGroupPartition( override val index: Int, val narrowDeps: Array[Option[NarrowCoGroupSplitDep]]) extends Partition with Serializable { override def hashCode(): Int = index override def equals(other: Any): Boolean = super.equals(other)}
2.這個NarrowCoGroupSplitDep的主要功能就是序列化,為了避免重複,對rdd做了瞬態註解
/** 這個NarrowCoGroupSplitDep的主要功能就是序列化,為了避免重複,對rdd做了瞬態註解*/private[spark] case class NarrowCoGroupSplitDep( @transient rdd: RDD[_], //瞬態的欄位不會被序列化,適用於臨時變數 @transient splitIndex: Int, var split: Partition ) extends Serializable { @throws(classOf[IOException]) private def writeObject(oos: ObjectOutputStream): Unit = Utils.tryOrIOException { // Update the reference to parent split at the time of task serialization split = rdd.partitions(splitIndex) oos.defaultWriteObject() }}
回到CoGroupedRDD上來,先看這個RDD的依賴是如何劃分的:
/* * 簡單看下CoGroupedRDD重寫的RDD的getDependencies: * 如果兩個rdd的分區函數相同就是窄依賴 * 否則就是寬依賴 */ override def getDependencies: Seq[Dependency[_]] = { rdds.map { rdd: RDD[_] => if (rdd.partitioner == Some(part)) { /*如果分區函數不為None 對應窄依賴*/ logDebug("Adding one-to-one dependency with " + rdd) new OneToOneDependency(rdd) } else { logDebug("Adding shuffle dependency with " + rdd) new ShuffleDependency[K, Any, CoGroupCombiner]( rdd.asInstanceOf[RDD[_ <: Product2[K, _]]], part, serializer) } } }
CoGroupedRDD.getPartitions 返回一個帶有Partitioner.numPartitions個分區類型為CoGroupPartition的數組
/* * 這裡返回一個帶有Partitioner.numPartitions個分區類型為CoGroupPartition的數組 */ override def getPartitions: Array[Partition] = { val array = new Array[Partition](part.numPartitions) for (i <- 0 until array.length) { // Each CoGroupPartition will have a dependency per contributing RDD //rdds.zipWithIndex 這個是產生一個(rdd,rddIndex)的索引值對,可以查看Seq或者Array的API //繼續跟到CoGroupPartition這個Partition,其是和Partition其實區別不到,只是多了一個變數narrowDeps //回來看NarrowCoGroupSplitDep的構造,就是傳入了每一個rdd和分區索引,以及分區,其可以將分區序列化 array(i) = new CoGroupPartition(i, rdds.zipWithIndex.map { case (rdd, j) => // Assume each RDD contributed a single dependency, and get it dependencies(j) match { case s: ShuffleDependency[_, _, _] => None case _ => Some(new NarrowCoGroupSplitDep(rdd, i, rdd.partitions(i))) } }.toArray) } array }
好,現在弱弱的總結下CoGroupedRDD,其類型大概是(k,(Array(CompactBuffer[v1]),Array(CompactBuffer[v2]))),這其中用到了內部的封裝,以及compute函數的實現
有興趣的可以繼續閱讀下源碼,這一部分就不介紹了。
下面還是幹點正事,把join運算元的整體簡單理一遍:
1.join 運算元內部使用了cogroup運算元,這個運算元返回的是(key,(v1,v2))這種形式的元組
2.深入cogroup運算元,發現其根據rdd1,rdd2建立了一個CoGroupedRDD
3.簡要的分析了CoGroupedRDD的依賴關係,看到如果兩個rdd的分區函數相同,那麼產生的rdd分區數不變,它們之間是一對一依賴,也就是窄依賴,從而可以減少依次shuffle
4. CoGroupedRDD的分區函數就是將兩個rdd的相同分區索引的分區合成一個新的分區,並且通過NarrowCoGroupSplitDep這個類實現了序列化
5.具體的合并過程還未記錄,之後希望可以補上這部分的內容
Spark join 源碼跟讀記錄