coalesce 与 repartition的用法与区别_rdd.coalesce(100)-程序员宅基地

它们两个都是RDD的分区进行重新划分，repartition只是coalesce接口中shuffle为true的简易实现
rdd 由多变少可以不进行shuffe，由少变多必须进行shuffe

先看coalesce：

 /**
   * Return a new RDD that is reduced into `numPartitions` partitions.
   *
   * This results in a narrow dependency, e.g. if you go from 1000 partitions
   * to 100 partitions, there will not be a shuffle, instead each of the 100
   * new partitions will claim 10 of the current partitions. If a larger number
   * of partitions is requested, it will stay at the current number of partitions.
   *
   * However, if you're doing a drastic coalesce, e.g. to numPartitions = 1,
   * this may result in your computation taking place on fewer nodes than
   * you like (e.g. one node in the case of numPartitions = 1). To avoid this,
   * you can pass shuffle = true. This will add a shuffle step, but means the
   * current upstream partitions will be executed in parallel (per whatever
   * the current partitioning is).
   *
   * @note With shuffle = true, you can actually coalesce to a larger number
   * of partitions. This is useful if you have a small number of partitions,
   * say 100, potentially with a few partitions being abnormally large. Calling
   * coalesce(1000, shuffle = true) will result in 1000 partitions with the
   * data distributed using a hash partitioner. The optional partition coalescer
   * passed in must be serializable.
   */

注释的大致意思就是假设父rdd 1000分区，然后调用coalesce(100)，实际上就是将父rdd的1000分区分成100组，每组10个，叫做partitionGroup，每个partitionGroup作为coalescedrdd的一个分区，在compute方法中迭代处理，以此来避免shuffle。

coalesce函数总共三个参数：分区数，是否进行shuffle(默认不shuffle)，Coalesce分区器(用来决定哪些父rdd的分区组成一组，作为一个partitiongroup，也即是决定了coalescedrdd的分区情况)。

def coalesce(numPartitions: Int, shuffle: Boolean = false,
               partitionCoalescer: Option[PartitionCoalescer] = Option.empty)
              (implicit ord: Ordering[T] = null)
      : RDD[T] = withScope {
    require(numPartitions > 0, s"Number of partitions ($numPartitions) must be positive.")
    if (shuffle) {
      /** Distributes elements evenly across output partitions, starting from a random partition. */
      val distributePartition = (index: Int, items: Iterator[T]) => {
        var position = new Random(hashing.byteswap32(index)).nextInt(numPartitions)
        items.map { t =>
          // Note that the hash code of the key will just be the key itself. The HashPartitioner
          // will mod it with the number of total partitions.
          position = position + 1
          (position, t)
        }
      } : Iterator[(Int, T)]
      // include a shuffle step so that our upstream tasks are still distributed
      new CoalescedRDD(
        new ShuffledRDD[Int, T, T](
          mapPartitionsWithIndexInternal(distributePartition, isOrderSensitive = true),
          new HashPartitioner(numPartitions)),
        numPartitions,
        partitionCoalescer).values
    } else {
      new CoalescedRDD(this, numPartitions, partitionCoalescer)
    }
  }

在看repartition

 def repartition(numPartitions: Int)(implicit ord: Ordering[T] = null): RDD[T] = withScope {
    coalesce(numPartitions, shuffle = true)
  }

本文链接：https://blog.csdn.net/weixin_38401971/article/details/105727612

原作者删帖不实内容删帖广告或垃圾文章投诉

智能推荐

Spring Boot 获取 bean 的 3 种方式！还有谁不会？，Java面试官_springboot2.7获取bean-程序员宅基地

文章浏览阅读1.2k次，点赞35次，收藏18次。AutowiredPostConstruct 注释用于在依赖关系注入完成之后需要执行的方法上，以执行任何初始化。此方法必须在将类放入服务之前调用。支持依赖关系注入的所有类都必须支持此注释。即使类没有请求注入任何资源，用 PostConstruct 注释的方法也必须被调用。只有一个方法可以用此注释进行注释。_springboot2.7获取bean

Logistic Regression Java程序_logisticregression java-程序员宅基地

文章浏览阅读2.1k次。理论介绍节点定义package logistic;public class Instance { public int label; public double[] x; public Instance(){} public Instance(int label,double[] x){ this.label = label; th_logisticregression java

linux文件误删除该如何恢复？，2024年最新Linux运维开发知识点-程序员宅基地

文章浏览阅读981次，点赞21次，收藏18次。本书是获得了很多读者好评的Linux经典畅销书**《Linux从入门到精通》的第2版**。下面我们来进行文件的恢复，执行下文中的lsof命令，在其返回结果中我们可以看到test-recovery.txt (deleted)被删除了，但是其存在一个进程tail使用它，tail进程的进程编号是1535。我们看到文件名为3的文件，就是我们刚刚“误删除”的文件，所以我们使用下面的cp命令把它恢复回去。命令进入该进程的文件目录下，1535是tail进程的进程id，这个文件目录里包含了若干该进程正在打开使用的文件。

流媒体协议之RTMP详解-程序员宅基地

文章浏览阅读10w+次，点赞12次，收藏72次。RTMP(Real Time Messaging Protocol)实时消息传输协议是Adobe公司提出得一种媒体流传输协议，其提供了一个双向得通道消息服务，意图在通信端之间传递带有时间信息得视频、音频和数据消息流，其通过对不同类型得消息分配不同得优先级，进而在网传能力限制下确定各种消息得传输次序。_rtmp

微型计算机2017年12月下,2017年12月计算机一级MSOffice考试习题(二)-程序员宅基地

文章浏览阅读64次。2017年12月的计算机等级考试将要来临!出国留学网为考生们整理了2017年12月计算机一级MSOffice考试习题，希望能帮到大家，想了解更多计算机等级考试消息，请关注我们，我们会第一时间更新。2017年12月计算机一级MSOffice考试习题(二)一、单选题1). 计算机最主要的工作特点是( )。A.存储程序与自动控制B.高速度与高精度C.可靠性与可用性D.有记忆能力正确答案：A答案解析：计算...

20210415web渗透学习之Mysqludf提权（二）（胃肠炎住院期间转）_the provided input file '/usr/share/metasploit-fra-程序员宅基地

文章浏览阅读356次。在学MYSQL的时候刚刚好看到了这个提权，很久之前用过别人现成的，但是一直时间没去细想，这次就自己复现学习下。 0x00 UDF 什么是UDF？ UDF (user defined function)，即用户自定义函数。是通过添加新函数，对MySQL的功能进行扩充，就像使..._the provided input file '/usr/share/metasploit-framework/data/exploits/mysql

随便推点

webService详细-程序员宅基地

文章浏览阅读3.1w次，点赞71次，收藏485次。webService一 WebService概述1.1 WebService是什么WebService是一种跨编程语言和跨操作系统平台的远程调用技术。Web service是一个平台独立的，低耦合的，自包含的、基于可编程的web的应用程序，可使用开放的XML（标准通用标记语言下的一个子集）标准...

Retrofit(2.0)入门小错误 -- Could not locate ResponseBody xxx Tried: * retrofit.BuiltInConverters_已添加addconverterfactory 但是 could not locate respons-程序员宅基地

文章浏览阅读1w次。前言照例给出官网：Retrofit官网其实大家学习的时候，完全可以按照官网Introduction，自己写一个例子来运行。但是百密一疏，官网可能忘记添加了一句非常重要的话，导致你可能出现如下错误：Could not locate ResponseBody converter错误信息：Caused by: java.lang.IllegalArgumentException: Could not l_已添加addconverterfactory 但是 could not locate responsebody converter

一套键鼠控制Windows+Linux——Synergy在Windows10和Ubuntu18.04共控的实践_linux 18.04 synergy-程序员宅基地

文章浏览阅读1k次。一套键鼠控制Windows+Linux——Synergy在Windows10和Ubuntu18.04共控的实践Synergy简介准备工作（重要）Windows服务端配置Ubuntu客户端配置配置开机启动Synergy简介Synergy能够通过IP地址实现一套键鼠对多系统、多终端进行控制，免去了对不同终端操作时频繁切换键鼠的麻烦，可跨平台使用，拥有Linux、MacOS、Windows多个版本。Synergy应用分服务端和客户端，服务端即主控端，Synergy会共享连接服务端的键鼠给客户端终端使用。本文_linux 18.04 synergy

nacos集成seata1.4.0注意事项_seata1.4.0 +nacos 集成-程序员宅基地

文章浏览阅读374次。写demo的时候遇到了很多问题，记录一下。安装nacos1.4.0配置mysql数据库，新建nacos_config数据库，并根据初始化脚本新建表，使配置从数据库读取，可单机模式启动也可以集群模式启动，启动时 ./start.sh -m standaloneapplication.properties 主要是db部分配置## Copyright 1999-2018 Alibaba Group Holding Ltd.## Licensed under the Apache License,_seata1.4.0 +nacos 集成

iperf3常用_iperf客户端指定ip地址-程序员宅基地

文章浏览阅读833次。iperf使用方法详解 iperf3是一款带宽测试工具，它支持调节各种参数，比如通信协议，数据包个数，发送持续时间，测试完会报告网络带宽，丢包率和其他参数。安装 sudo apt-get install iperf3 iPerf3常用的参数： -c ：指定客户端模式。例如：iperf3 -c 192.168.1.100。这将使用客户端模式连接到IP地址为192.16..._iperf客户端指定ip地址

浮点性(float)转化为字符串类型自定义实现和深入探讨C++内部实现方法_c++浮点数转字符串精度损失最小-程序员宅基地

文章浏览阅读7.4k次。写这个函数目的不是为了和C/C++库中的函数在性能和安全性上一比高低，只是为了给那些喜欢探讨函数内部实现的网友，提供一种从浮点性到字符串转换的一种途径。浮点数是有精度限制的，所以即使我们在使用C/C++中的sprintf或者cout 限制，当然这个精度限制是可以修改的。比方在C++中，我们可以cout.precision(10),不过这样设置的整个输出字符长度为10，而不是特定的小数点后1_c++浮点数转字符串精度损失最小