Casa Parallelization meeting minutes

Monday June 18th Room 317, 3:30pm MDT

Attendees:

Targets/tasks for June 18th - 25th

  • Complete installation of test hardware James
  • Complete or nearly complete test script James
  • Begin completion of pclean vs clean parametrization Kumar
  • Identify pcube ALMA tester Brian
  • Identify large number channel data set Butler

Discussion/Agenda

  • Benchmarking outline for final cluster orders
    • Use 400GB 3c147 data set, large A array image with 1K channels
      • Sufficiently large to measure multinode performance with pcube and pclean
      • Re-demonstrate non-imaging vs imaging task times and I/O performance
      • Effects of per core and per node parallelization per task
    • Performance against variety of hardware
      • 5500, 5600 Westmere Xeon 6core processorss, E5-2400 and E5-2600 Sandy Bridge Xeon 6 and 8 core processors
        • Generate various imaging/cost measures (e.g. visibilities imaged per dollar) (assumption is large number of cheapest processors is best)
        • Measure parallelization performance and impacts due to cache starvation
      • Dell, SGI and Supermicro motherboards with similar processors and memory
        • Measure Infinband and over all performance and cost
    • Demonstrate memory impacts for various sized images, particularly swapping impact on performance
      • Need better description of what imaging tasks and sizes must be supported
    • Ultimately generate optimal price configuration of processor vs memory tradeoffs vs specific imaging cases to create hetergenous cluster config
  • Discussion of priorities for 3.5 imaging
    • Full parameter implementation of clean to pclean (e.g. multifield, utilitarian options) versus memory consumption issues
      • Devote 2 weeks to pclean implementation
      • Debate about examine memory vs scratchless/cal library interface

End of discussion due to time

  • Memory requirements in imaging
    • Total memory scaling is generally understood
    • Kumar has identified expected vs realized memory usage discrepency.
    • Need to understand better how temp image selection works
  • Describe and implement engine vs threaded gridder determination as a function of memory
    • Need dynamic image partitioning scheme for 2, 3, 4 or more threads.
    • (Total memory - some delta ) / memory per engine = number of engines
    • int (Total cores / number engines) = number of threads per engine
  • Update on Partitioning status (null selections, non-conforming tasks) Jeff
    • First tasks being implemented (flagger, calibration), some read in parallel, all can read serially, writers are still problematic
    • Re-examination of Parallel-go for cluster specification (cores vs engines, memory)
  • Model data tests James
    • Pending completion of test script
    • James to get with Kumar in late June
  • Memory resident subtable tests James late June
    • Pending completion of test script
    • Check that it affects number of FDs as well
  • Spectral imaging tests James
    • Understand channel chunk choice effect on performance James
    • Pending completion of test script

CASA HPC Initiatives for the 3.5 Cycle

  • Tentative list from last meeting
    • Reduce in memory requirements during imaging (maybe removed)
    • Parallelize multi-field imaging
    • Parallelize MFS terms > 1 in the gridding step
    • Parallelize Cube imaging scatter gather
    • Expand threaded gridder case to >4 threads
    • Implement core affinity for threaded gridder case (first out)
    • Parallel framework (partitioning, parallel go) work in ESO
    • Filler to write a MMS in parallel

Deferred Topics

ARDG Targets

-- JamesRobnett - 2012-06-25
Topic revision: r1 - 2012-06-25, JamesRobnett
This site is powered by FoswikiCopyright © by the contributing authors. All material on this collaboration platform is the property of the contributing authors.
Ideas, requests, problems regarding NRAO Public Wiki? Send feedback