Showing posts with label Scale-Out File Server. Show all posts
Showing posts with label Scale-Out File Server. Show all posts

Monday, 23 July 2018

Mellanox SwitchX-2 and Spectrum OS Update Grids

We're in the process of building out a new all-flash based Kepler-64 2-node cluster that will be running the Scale-Out File Server Role. This round of testing will have several different rounds to it:

  1. Flat Flash Intel SSD DC S4500 Series
    • All Intel SSD DC S4500 Series SATA SSD x24
  2. Intel NVMe PCIe AIC Cache + Intel SSD DC S4500
    • Intel NVMe PCIe AIC x4
    • Intel SSD DC S4500 Series SATA SSD x24
  3. Intel Optane PCIe AIC + Intel SSD DC S4500
    1. Intel Optane PCIe AIC x4
    2. Intel SSD DC S4500 Series SATA SSD x24

Prior to running the above tests we need to update the operating system on our two Mellanox SwitchX-2 MSX1012B series switches as we've been quite busy with other things!

Their current OS level is 3.6.4006 so just a tad bit out of date.

image

The current OS level for SwitchX-2 PPC switches is 3.6.8008. And, as per the Release Notes for this OS version we need to do a bit of a Texas Two-Step to get our way up to current.

image

image

Now, here's the kicker: There is no 3.6.5000 on Mellanox's download site. The closest version to that is 3.6.5009 which provides a clarification on the above:

image

Okay, so that gets us to 3.6.5009 that in turn gets us to 3.6.6106:

image

And that finally gets us to 3.6.8008:

image

Update Texas Two Step

To sum things up we need the following images:

  1. 3.6.4122
  2. 3.6.5009
  3. 3.6.6106
  4. 3.6.8008

Then, it's a matter of time and a bit of patience to run through each step as the switches can take a bit of time to update.

image

A quick way to back up the configuration is to click on the Setup button then Configurations then click the initial link.

image

Copy and paste the output into a TXT file as it can be used to reconfigure the switch if need-be via the Execute CLI commands window just below it.

As always, it pays to read that manual eh! ;)

NOTE: Acronym Finder: AIC = Add-in Card so not U.2.

Oh, and be patient with the downloads as they are _slow_ as molasses in December as of this writing. :(

image

Philip Elder
Microsoft High Availability MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
www.s2d.rocks !
Our Web Site
Our Cloud Service

Friday, 3 November 2017

A Little Plug for Mellanox and RoCE RDMA

RoCE (RDMA over Converged Ethernet) via Mellanox NICs and switches is our primary fabric choice for Storage Spaces Direct (S2D) and Scale-Out File Server (SOFS) to Hyper-V compute cluster fabric.

With the Mellanox MSX1012X 10GbE switch we can deploy a pair of them along with a pair of ConnectX-4 Lx dual port NICs per node for about the same cost as a pair of NETGEAR XS716T 10GbE switches and a pair of Intel X540/X550-T2 10GbE RJ45 based NICs per node.

We have a great business relationship with Mellanox. They are great folks to work with and their product support is second to none.

I was honoured to be asked to use a portion of my presentation for MVPDays to create the following video that is resident on Mellanox's YouTube channel.

Hopefully the video comes out okay as embedding it was a bit of a chore.

Thanks for reading and have a great weekend!

Philip Elder
Microsoft High Availability MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Our Cloud Service
Twitter: @MPECSInc

Tuesday, 16 February 2016

Cluster 101: Some Hyper-V and SOFS Cluster Basics

Our focus here at MPECS Inc. has grown into providing cluster-based solutions to clients near and far over the last eight years or so as well as cluster infrastructure solutions for small to medium I.T. shops.

There were so many misconceptions when we started the process to build out our first Hyper-V cluster in 2008.

The call in to us was for a large food manufacturing company that had a very specific requirement for their SQL, ColdFusion, and mail workloads to be available. The platform of choice was the Intel Modular Server with an attached Promise VTrak E310sD for extra storage.

So, off we went.

We procured all of the hardware through the Intel and Promise demo program. There was _no_ way we were going to purchase close to $100K of hardware on our own!

Back then, there was a dearth of documentation … though that hasn’t changed all that much! ;)

It took six months of trial and error plus working with the Intel, Promise, LSI, and finally contacts at Microsoft to figure out the right recipe for standing up a Hyper-V cluster.

Once we had everything down we deployed the Intel Modular Server with three nodes and the Promise VTrak E310sD for extra storage.

Node Failure

One of the first discoveries: A cluster setup does not mean the workload stays up if the node it’s on spontaneously combusts!

What does that mean? It means that when a node suddenly goes offline because of a hardware failure the guest virtual machines get moved over to an available node in a powered off state.

To the guest OS it is as if someone hit the reset button on the front of a physical server. And, as anyone that has experienced a failed node knows the first prompt when logging in to the VM is the “What caused the spontaneous restart” prompt.

Shared Storage

Every node in a Hyper-V cluster needs identical access to the storage the VHD(x) files are going to reside on.

In the early days, there really was not a lot of information indicating exactly what this meant. Especially since we decided right from day one to avoid any possible solution set based on iSCSI. Direct Attached Storage (DAS) via SAS was the way we were going to run with. The bandwidth was vastly superior with virtually no latency. No other shared storage in a cluster setting could match the numbers. And, to this day the other options still can’t match DAS based SAS solutions.

It took some time to figure out, but in the end we needed a Shared Storage License (SharedLUNKey) for the Intel Modular Server setup and a storage shelf with the needed LUN Sharing and/or LUN Masking plus LUN Sharing depending on our needs.

We had our first Hyper-V cluster!

Storage Spaces

When Storage Spaces came along in 2012 RTM we decided to venture into Clustered Storage Spaces via 2 nodes and a shared JBOD. That process took about two to three months to figure out.

Our least expensive cluster option based on this setup (blog post) is deployed at a 15 seat accounting firm. The cost versus the highly available workloads benefit ratio is really attractive. :)

We have also ventured into providing backend storage via Scale-Out File Server clusters for Hyper-V cluster frontends. Fabric between the two starts with 10GbE and SMB Multichannel.

Networking

All Broadcom and vendor rebranded Broadcom NICs require VMQ disabled for each network port!

A best practice for setting up each node is to have a minimum of four ports available. Two for the management network and Live Migration network and two for the virtual switch team. Our preference is for a pair of Intel Server Adapter i350-T4s set up as follows:

  • Port 0: Management Team (both NICs)
  • Port 1 and 2: vSwitch (no host OS access both NICs)
  • Port 3: Live Migration networks (LM0 and LM1)

For higher end setups, we install at least one Intel Server Adapter X540-T2 to bind our Live Migration network to each port. In a two node clustered Storage Spaces setting the 10GbE ports are direct connected.

Enabling Jumbo Frames is mandatory for any network switch and NIC carrying storage I/O or Live Migration.

Hardware

In our experience GHz is king over cores.

The maximum amount of memory per socket/NUMA node that can be afforded should be installed.

All components that can be should be run in pairs to eliminate as many single points of failure (SPFs) as is possible.

  • Two NICs for the networking setup
  • Two 10GbE NICs at the minimum for storage access (Hyper-V <—> SOFS),
  • Two SAS HBAs per SOFS node
  • Two power supplies per node

On the Scale-Out File Server cluster and Clustered Storage Spaces side of things one could scale up the number of JBODs to provide enclosure resilience thus protecting against a failed JBOD.

The new DataON DNS-2670 70-bay JBOD supports eight SAS ports per controller for a total of 16 SAS ports. This would allow us to scale out to eight SOFS nodes and eight JBODs using two pairs of LSI 9300-16e (PCIe 8x)  or the higher performance LSI 9302-16e (PCIe 16x) SAS HBAs per node! Would we do it? Probably not. Three or four SOFS nodes would be more than enough to serve the eight direct attached JBODs. ;)

Know Your Workloads

And finally, _know your workloads_!

Never, ever, rely on a vendor for performance data on their LoB or database backend. Always make a point of watching, scanning, and testing an already in-place solution set for performance metrics or the lack thereof. And, once baselines have been established in testing the results remain private to us.

The two key ingredients in any standalone or cluster virtualization setting are:

  1. IOPS
  2. Threads
  3. Memory

A balance must be struck between those three relative to the budget involved. It is our job to make sure our solution meets the workload requirements that have been placed before us.

Conclusion

We’ve seen a progression in the technologies we are using to deploy highly available virtualization and storage solutions.

While the technology does indeed change over time the above guidelines have stuck with us since the beginning.

Philip Elder
Microsoft High Availability MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Thursday, 28 January 2016

Cluster: A Simple Cluster Storage Setup Guide

In a cluster setting we have a set way to configure our shared storage whether it resides on a SOFS (Scale-Out File Server) cluster or some sort of network based storage.

First, the process to set up the storage itself:

  1. Configure the LUN
    • LUN ID must be identical for all Hyper-V nodes for SAN/NAS
  2. Connect all nodes to the storage
    • iSCSI Target for SAN/NAS
  3. Format NTFS and set OFFLINE on Node01
  4. Node2 and up ignore Initialize in Disk Management and set OFFLINE
    • This step is optional depending on the setup

When it comes to the storage we configure the following LUNs for all of our cluster setups;

  1. 1.5GB LUN
    • Set up for the Witness Disk
    • Add to Cluster Storage but NOT CSV
  2. ???GB LUN
    • Sum of all physical RAM on the nodes plus 150GB
    • Add to Cluster Shared Volumes
    • All Hyper-V nodes set to deliver VM settings files to this location
    • Don’t forget that Hyper-V writes a file that is equivalent in size for _all_ VMs running on the cluster or standalone host!
  3. Minimum 50% Storage LUN x2
    • Divide the remaining storage into two or more LUNs depending on workload and storage requirements
    • A minimum of 2 LUNs allows for storage load to be shared across the SAN’s two storage controllers, the two iSCSI networks, and the two or more Hyper-V nodes

In a SOFS setting we set up a File Share Witness for our Hyper-V compute clusters and deliver the HA shares via SMB Multichannel and a minimum of 10GbE for the VHDX files.

PowerShell

The PowerShell steps for any of the above are here to avoid copy and paste issues.

Set Default Paths:

Set-VMHost -VirtualHardDiskPath “C:\ClusterStorage” –VirtualMachinePath “C:\ClusterStorage\Volume1”

We point the VHDX setting to the CSV root just in case. Our PowerShell scripts for setting up VMs put the VHDX files into the right storage location.

Set Quorum Up:

Set-ClusterQuorum -NodeAndDiskMajority "Cluster Virtual Disk (Witness Disk)"

Philip Elder
Microsoft High Availability MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Wednesday, 17 June 2015

What's up, what's been happing, and what will be happening.

Wow, it's been a while hasn't it? :)

We've been _very_ busy with our business as well as a Cloud services start-up and Third Tier is keeping me hopping too.

I have a regular monthly Webinar via Third Tier where we've been spending time on the Third Tier product called "Be the Cloud". It is a solution set developed to provide a highly available backend for client facing services based on our SBS (Small Business Solution).

We, that is my family, took a much needed break in May for a couple of weeks of downtime as we'd not had any pause for a good 18 months prior. We were ready for that.

So, why the blogging pause?

There are a number of reasons.

One is that I've been so busy researching and working on new things that there hasn't been a lot of time left over for writing them all out. Ongoing client needs are obviously a part of that too.

Another had to do with waiting until we were okay to publish information on the upcoming Windows Server release. We Cluster MVPs, and others, were privileged to be very deeply involved with the early stages of the new product. But, we were required to remain mum. So, instead of risking anything I decided to hold off on publishing anything Server vNext related.

Plus, we really didn't have a lot of new content to post since we've about covered the gamut in Windows Server 2012 RTM/R2 and Windows Desktop. Things have been stable on that front other than a few patch related bumps in the road. So, nothing new there meant nothing new to write about. ;)

And finally, the old grey matter just needed a break. After all, I've been writing on this blog since the beginning of 2007! :)

So, what does this mean going forward?

It means that we will begin publishing content on a regular basis again once we've began serious work with Windows Server vNext.

We have a whole host of lab hardware on the way that has a lot to do with what's happening in the new version of Windows Server that ties into our v2 for Be the Cloud and our own Cloud services backend.

We're also establishing some new key vendor relationships that will broaden our solution matrix with some really neat new features. As always, we build our solution sets and test them vigorously before considering a sale to a client.

And finally, we're reworking our PowerShell library into a nice and tidy OneNote notebook set to help us keep consistent across the board. This is quite time consuming as it becomes readily apparent that many steps are in the grey matter but not in Notepad or OneNote.

Things we're really excited about:
  • Storage Spaces Direct (S2D)
  • Storage Replication
  • Getting the Start Menu back for RDSH Deployments
    • Our first deployment on TP2 is going to happen soon so hopefully we do indeed have control over that feature again!
  • Deploying our first TP2 RDS Farm
  • Intel Server Systems are on the S2D approval list!
    • The Intel Server System R2224WTTYS is an excellent platform
  • Promise Storage J5000 series JBODs just got on the Storage Spaces approved list.
    • We've had a long history with Promise and are looking forward to re-establishing that relationship.
  • We've started working with Mellanox for assistance with RDMA Direct and RoCE.
  • 12Gb SAS in HBAs and JBODs rocks for storage
    • 2 Node SOFS Cluster with JBOD is 96Gbps of aggregate ultra-low latency SAS bandwidth per node!
  • NVMe based storage (PCIe direct)
The list could go on and on as they come to mind. :)

Thank you all for your patience with the lack of posting lately. And, thank you all for your feedback and support over the years. It has been a privilege to get to know some of you and work with some of you as well.

We are most certainly looking forward to the many things we have coming down the pipe. 2015 is shaping up to be our best year ever with 2016 looking to build on that!

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Friday, 12 September 2014

MANDATORY: Intel JBOD2224S2DP Firmware Update for Same Enclosure ID Storage Spaces Problem

Intel has released a new firmware for the Intel JBOD2224S2DP storage enclosure that deals with the enclosure delivering the same Enclosure ID to Windows Server 2012 R2 Storage Spaces.

Why is this firmware mandatory?

Because up until now when two or more Intel JBOD2224S2DP units were connected to Scale-Out File Server nodes and one ran the Get-StorageEnclosure PowerShell command one would get the same ID back for every one.

The firmware problem killed Storage Spaces enclosure resilience. What is that you ask?

In Storage Spaces, with a Two-Way or Three-Way mirror one can have three enclosures set up to allow one to drop out completely and things keep going.

If one has configured five enclosures with a Three-Way Mirror then the Storage Spaces setup can tolerate two enclosures dropping out.

If the Intel JBOD unit is already in a production setting with plans to add more enclosures at a later date then it is important to note that this firmware update would be required prior to adding the new units.

For new setups, this firmware update should be a part of the preparation steps for the JBOD prior to implementation or baseline testing.

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Thursday, 26 June 2014

CRITICAL: Seagate 1200 SSD Firmware Update Required for 2012 R2 Storage Spaces

We got hit with this today:

image

  • Get-PhysicalDisk

Since this was our second run at standing up this Scale-Out File Server cluster with things not working as expected we began to dig in.

During a Space creation process the symptom of the Volume Format phase hitting an error happened. We jumped into PowerShell to poll the disks right then and saw the above.

The Disks:

image

  • Get-PhysicalDisk | where MediaType -eq SSD | ft Model,FirmwareVersion -AutoSize

The PowerShell to get the above information along with the serial numbers:

  • Get-PhysicalDisk | where MediaType -eq SSD | ft Model,FirmwareVersion,SerialNumber -AutoSize

Go to Seagate's Support site and choose the Download Finder.

image

Enter the serial number, just the first set of digits before the four 0000 pattern as the number was repeated twice for us, your country, and then under Certificate click on the Click here link and _not_ the Email Me link.

image

The highlighted link downloads the actual firmware ZIP file.

A copy of Seagate's SeaTools is required to update the drive's firmware.

If the SSDs are in a cluster setting, as they are here, make sure to properly drain the nodes and shut down the cluster (TechNet). Then shut down all nodes but one.

Run the firmware update and reboot that node. Bring the Cluster back online and fire up the other nodes one by one.

NOTE: We have _not_ tested this firmware yet.

A Microsoft Forum post brought us into the right direction: Clustered Storage space degraded and SSD disks "starting".

Remember how it has been mentioned ad nauseam on this blog about how we are very careful about testing our deployments? Well, in this case we were a part of the planning phases for this cluster but did not have access to these particular SSDs prior to deploying in a Michigan Data Centre.

This situation sure brings home the point that we _always_ need to test our setups before deploying them at client sites or on behalf of our clients.

EDIT: Brain was five steps ahead of fingers so the "Firmware Update" in the title never made it into the original post! :)

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Thursday, 19 June 2014

Cluster Node BIOS and Hardware Configuration Tips

Here are some tips for configuring the nodes in a Hyper-V or Scale-Out File Server failover cluster.

Staggered Start

image

  • Stagger the node start times to give storage enough time to come online

One of the important tests to run when working with a new JBOD unit or storage shelf is to time the unit's power-up to production ready time.

In the case of the Intel JBOD2224S2DP with 24 Seagate Savvio spindles installed the staggered start of each disk group actually takes a bit of time to process. So, we set our Grizzly Pass servers to start at 150 seconds and up for each storage node and then 210 and up for each Hyper-V server node.

Processor C States

image

  • Processor C States are set to Disabled

Why the C States interfere with storage access and transfer abilities is a bit of a mystery but they do need to be turned off.

Also, take careful notes of all BIOS settings set up on one node and make sure to set all other node's BIOS settings to the same ones.

Performance Setting

Pedal to the metal:

image

Make sure the performance profiles are set to maximum!

We need all available power at all times.

PXE Boot

We suggest turning PXE Boot and the NIC's option ROM off.

image

Confirm in the Boot Order manager that there are no NICs available for boot. If any show up there make sure to disable them.

While in the NIC configuration settings one can make a note of the NIC MAC addresses to help with configuration further on into the node setup process.

Reboot and OS Boot Checks

We've seen some issues with OS Boot Watchdog Timers:

image

Most modern BIOS firmware should be able to sense that Windows Server 2012 R2 has booted and settled into its working role. But, we have seen cases in older BIOS versions where the server would mysteriously reboot after 10 minutes (we timed it after noticing that the reboots were happening close to the same time).

Boot Options

And finally, for now we are not enabling EFI Optimized Boot options on our nodes:

image

We need to run some tests with 2012 R2 U1 before we commit to the new setup in production.

Make sure to disable the USB Boot Priority or that OS Load flash drive will be booted to on node reboots!

Cluster Node Configuration

When it comes to setting up the node specifications one needs to choose carefully.

This again is one area where Intel Server Systems outshine Tier 1.

The Intel Server System R1208JP4OC is an Intel Xeon Processor E5-2600 v1/v2 series 1U server with a single socket. The big plus to this server is the ability to have two SAS HBAs and two 10GbE or 56Gb InfiniBand cards installed.

As far as we know no Tier 1 single socket 1U server shares this ability anywhere. So, we get a really good performing server at an excellent entry level price point.

We make sure to design our clusters around their intended purpose at the storage, Scale-Out File Server, and Hyper-V levels.

With these tools that are included in Windows Server 2012 RTM/R2 we have an amazing ability to build a single asymmetric cluster (2 nodes and 1 JBOD) at a very lucrative price that fits in really well at the SMB level (12-13 seats plus - yes, we sell clusters into SMB) right up to a million IOPS plus transaction oriented cluster.

Remember that consistency in hardware, firmware, settings, and drivers is the key to cluster performance and stability.

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Wednesday, 18 June 2014

SOFS, Storage Spaces, and a Big Thanks to the Intel Technology Provider Program!

What was once the Intel Channel Program and now ITP has been very generous to us over the years.

We make no bones about our support of both the program but also the excellent Intel Server Systems and Intel Storage Systems that we deploy on a regular basis.

With the introduction of the Grizzly Pass product line we received a product that was bang-on with Dell, HP, and IBM feature for feature, construction quality for construction quality, with two very significant advantages to the Intel product:

  1. Flexibility
    • We can utilize an extensive tested hardware list to custom configure our server and storage systems to order way beyond what Tier 1 offers even in their Build-to-Order programs.
    • We are able tune our configurations to very specific performance needs.
  2. Support
    • The folks on the other end of the support line are second to none. Some of the folks we have worked with have been our contact for cases over the last ten years or more! These folks know their stuff.
    • Advanced no questions asked warranty replacement for almost all products is also a huge asset.

This is the product stack we have been working on lately for our Proof-of-Concept testing for Scale-Out File Server failover clusters, Hyper-V over SMB via 10GbE provided for by two NETGEAR XS712T 10GbE switches, and Storage Spaces performance testing.

image

The top two servers are Intel R1208JP4OC 1U single socket servers supporting the Intel Xeon Processor E5-2600 v1/v2 series CPUs. They have dual Intel X540T2 NICs via I/O Module and PCIe add-in card along with a pair of Intel RS25GB008 SAS HBAs to provide connectivity to the Intel JBODs at the bottom.

Two of the Intel Server System R2208GZ4GC 2U dual socket servers were here for the last couple of months on loan from the Intel Technology Provider program. We have been using them extensively in our SOFS and Storage Spaces testing along with the other four servers that are our own.

One of the Intel Storage System JBOD2224S2DP units in the above picture is a seed unit provided to us by ITP as we are planning on utilizing this unit for our Data Centre deployments. The other two were purchased through Canadian distribution. Currently two are in a dedicated use configuration with the third to be used to test enclosure resilience in a Storage Spaces 3-Way Mirror configuration.

We have been acquiring HGST SAS SSDs in the form of first and second generation units with an aim to get into 12Gb SAS at some point down the road. We still have a few more first and second generation SSDs to go to reach our goal of 24 units total.

The second JBOD has 24 Seagate Savvio 10K SAS spindles that will be worked on in our next round of testing.

Our current HGST SAS SSD based IOPS testing average is about 375K on an 8 SSD disk set up in a Storage Spaces Simple configuration (similar to RAID 0):

image

We have designs on the board for providing enclosure resilient solutions that run into the millions of IOPS. As we move through our PoC testing we will continue to publish our results here.

We are currently working with Iometer for our baseline and PoC testing. SQLIO will also be utilized once we get comfortable with performance behaviours in our storage setups to fine tune things for SQL deployments.

Again, thanks to Scott P. and the Intel Technology Program for all of your assistance over the years. It is greatly appreciated. :)

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Monday, 9 June 2014

Storage Configuration: Know Your Workloads for IOPS or Throughput

Here we have a practical example of how devastating a poorly configured disk subsystem can be.
image
The above was one of the first Iometer test runs we did on our Storage Spaces setup. The above 45K IOPS was running on 17, yes seventeen, 100GB SSD400S.a HGST SAS SSDs.
Obviously the configuration was just whacked. :(
Imagine the surprise and disappointment one would have supplying a $100K SAN and ending up with the above results after the unit was put into production and the client was complaining that things were not happening anywhere near as fast as expected.
What we are discovering is that tuning a storage subsystem is an art.
There are so many factors that one needs to keep in mind as far as the types of workloads that will be running on the disk subsystem right through to the hardware driving it all.
After running a large number of tests using Iometer, and with some significant input from fellow MVP Tim Barrett, we are beginning to gain some insight into how to configure things for the given workload.
This is a snip taken of a Simple Storage Space utilizing _just two_ 100GB HGST SSD400S.a SAS SSDs (same disks as above):
image
Note how we are now running at 56K IOPS. :)
Microsoft has an awesome, in-depth, document on setting things up for Storage Spaces performance here:
We suggest firing the above article into OneNote for later reference as it will prove invaluable in figuring out the basics for configuring a Storage Spaces disk subsystem. It can actually provide a good frame of reference for storage performance in general.
Our goal for our Proof-of-Concept testing that we are doing was around 1M IOPS.
Given what we are seeing so far we will hopefully end up running at about 650K to 750K IOPS! That's not too shabby for our "commodity hardware" setup. :)
Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Thursday, 5 June 2014

Cluster Starter Labs: Hyper-V, Storage Spaces, and Scale-Out File Server

The following are a few ways to go about setting up a lab environment to test out various Hyper-V and Scale-Out File Server Clusters that utilize Storage Spaces to tie in the storage.

Asymmetric Hyper-V Cluster

  • (2) Hyper-V Nodes with single SAS HBA
  • (1) Dual Port SAS JBOD (must support SES-3)
In the above configuration we set up the node OS Roles and then enable Cluster. Once cluster is enabled we can import our not initialized shared storage into Cluster Disks and them move them over to Cluster Shared Volumes.
In this scenario one should split the storage up three ways.
  1. 1GB-2GB for Witness Disk
  2. 49.9% CSV 0
  3. 49.9% CSV 1
Once the virtual disks have been set up in Storage Spaces we run the quorum configuration wizard to set the witness disk up.
We use two CSVs in this setup so as to assign 50% of the available storage to each node. This shares the I/O load. Keep this in mind when looking to deploy this type of cluster into a client setting as well as the need to make sure all paths between the nodes and the disks are redundant (dual SAS HBAs and a dual expander/controller JBOD).

Symmetric Hyper-V Cluster with Scale-Out File Services

  • (2) Scale-Out File Server Nodes with single SAS HBA
  • (1) Dual Port SAS JBOD
  • (2) Hyper-V Nodes
For this particular set up we configure our two storage nodes in a SOFS cluster and utilize Storage Spaces to deliver our shares for Hyper-V to access. We will have a witness share for the Hyper-V cluster and then at least one file share for our VHDX files depending on how our storage is set up.

Lab Hardware

The HP MicroServer would be one option for server nodes. Dell C1100 1U off-lease servers can be found on eBay for a song. Intel RS25GB008 or LSI 6Gb SAS Host Bus Adapters (HBAs) are also easily found.
For the JBOD one needs to make sure the unit supports the full compliment of SAS commands being passed through to the disks. To run with cluster two SAS ports that access all of the storage installed in the drive bays is mandatory.
The Intel JBOD2224S2DP (WSC SS Site) is an excellent unit to work with that compares feature wise with DataON, Quanta, and the Dell JBODs now on the Windows Server Catalogue Storage Spaces List.
Some HGST UltraStar 100GB and 200GB SAS SSDs (SSD400 A and B Series) can be had via eBay every once in a while for SSD Tier and SSD Cache testing in Storage Spaces. We are running with the HGST product because it is a collaborative effort between Intel and HGST.

Storage Testing

For storage in the lab it is preferred to have at least 6 of the drives one would be using in production. With six drives we can run the following tests:
  • Single Drive IOPS and Throughput tests
    • Storage Spaces Simple
  • Dual Drive IOPS and Throughput tests
    • Storage Spaces Simple and Two-Way Mirror
  • Three Drive IOPS and Throughput tests
    • Storage Spaces Simple, Two-Way Mirror, and Three-Way Mirror
  • ETC to 6 drives+
There are a number of factors involved in storage testing. The main thing is to establish a baseline performance metric based on a single drive of each type.
A really good, and in-depth, read on Storage Spaces performance:
And, the Microsoft Word document outlining the setup and the Iometer settings Microsoft used to achieve their impressive 1M IOPS Storage Spaces performance:
Our previous blog post on a lab setup with a few suggested hardware pieces:
Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book
Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Friday, 23 May 2014

Three Intel Server Systems based Hyper-V and Scale-Out File Server Clusters

Here are three base Intel Server Systems configurations we are working on for our Intel Modular Server replacement in a Data Centre or client setting.

Unfortunately, the Intel JBOD does not self-power at this time. So, for SMB/SME solutions we will be supplying a DataON DNS-1640 2U JBOD as it will automatically power-up after a full power outage.

All solution sets are based on Windows Server 2012 R2 as a starting point for Hyper-V, Storage Spaces, and SOFS.

  • Option 1: Asymmetric Hyper-V Cluster via Storage Spaces CSV
    • Intel Server System R2208GZ4GC, Dual E5-2640, 128GB ECC or 256GB ECC, 120GB SSD RAID 1, dual SAS HBAs, add-in Intel i350T4 PCIe
    • Intel JBOD2224S2DP
  • Option 2: Hyper-V Cluster via SMBv3 Scale-Out File Server cluster and Storage Spaces
    • Intel Server System R1208JP4OC, E5-2640, 128GB ECC, 120GB SSD RAID 1, dual SAS HBAs, Intel X540T2 I/O Module, Intel X540T2 PCIe
    • Intel JBOD2224S2DP
    • Intel Server System R1208JP4OC, E5-2640, 128GB ECC, 120GB SSD RAID 1, Intel i350T4 PCIe, Intel X540T2 I/O Module, Intel X540T2 PCIe
    • NETGEAR XS712T 10GbE Switches
  • Option 3: Hyper-V Cluster via SMBv3 Scale-Out File Server cluster and Storage Spaces with enclosure resilience
    • (3) Intel Server System R2208GZ4GC, Dual E5-2640, 128GB ECC, 120GB SSD RAID 1, SIX SAS HBAs, Intel X540T2 I/O Module, Intel X540T2 PCIe
    • (3) Intel JBOD2224S2DP
    • (2) Intel Server System R2208GZ4GC, Dual E5-2640, 128GB ECC, 120GB SSD RAID 1, Intel i350T4 PCIe, Intel X540T2 I/O Module, Intel X540T2 PCIe
    • (2) NETGEAR 24-Port 10GbE Switches
  • Storage Networking Option
    • Option 2 and Option 3 can be facilitated by InfiniBand NICs and Switches
      • Enables RDMA and 56Gbps per connection
      • Microsoft's 1.4M IOPS demo based on InfiniBand backend
      • Intel Server Systems have an InfiniBand I/O Module with the second being a Mellanox PCIe

The first setup is relatively simple while the second two require some structuring around how the networking is configured to allow for SMB Multi-Channel on the storage network side.

At this point the above setups utilizing Intel Server Systems provide us with an amazing value for our IT budgets.

5 year warranties and next business day on-site support options can be had too.

We purchase our Intel Channel product primarily through ASI Canada. Ingram Micro, Synnex Canada, and Tech Data Canada are also Intel Authorized Distributors.

As an FYI we continue to build our own server systems because the experience proves to be invaluable when it comes to troubleshooting problems especially when software vendors are pointing fingers.

Building our own systems also gives us a very strong foundation for creating server configurations that will work with a client workload set.

And finally, it allows us to be very particular with Tier 1 vendors when it comes to creating a server configuration using their hardware.

EDIT: Note that we _always_ install a physical DC on our cluster networks. For option 1 it would probably be an HP MicroServer while the others would be a 1U single socket with some storage for ISOs.

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Monday, 21 April 2014

A Microsoft Cluster Troubleshooting Guide

Here are some of the tools we can use when troubleshooting a cluster, Scale-Out File Server, Hyper-V, and other cluster issues:

Failover Cluster Manager 

The FCM gives us the ability to dig into the various Windows Logs and delimit them by time, node, and log type.

  • FCM --> Cluster Name --> Cluster Events --> Query

image

We set up a few different queries out of the box. One with everything Cluster, Failover Clustering, and Hyper-V related. We then create a subset of queries. The various queries get saved to a local folder on the management DC/RSAT system.

Get-ClusterLog

The Get-ClusterLog PowerShell commandlet allows us to pull the full log set from one or all nodes. Note that the default output folder is \\NODE\C$\Windows\Cluster\Cluster.LOG (C:\Windows\Cluster\Cluster.LOG) unless specified in the command.

This log can be very busy and a bit of a challenge to work through. If one has a good idea of what to look for then the log can be quite informative.

  • Get-ClusterLog -Destination .
    • Places the log in the local directory (we create C:\Temp on all nodes for this kind of thing)
  • (get-cluster).ClusterLogLevel=5
    • There are five levels with 5 being the most verbose. Default level is 3 and best left there unless absolutely needed. Level 5 file can be large.

EDIT: The Default cluster log location is C:\Windows\Cluster\Reports\Cluster.log

Microsoft Message Analyzer

This is an in-depth tool. There is no way around it. Thus, a learning curve is required.

However, there is an amazing amount of information that we can then have at our fingertips and not only that colour coded!

image

image

One can use a series of filters under the log file settings to delimit by time period among others.

image

We can set up our columns:

image

Once we have our Cluster column, for example when looking for a problematic cluster component, we can set up a filter:

image

And that is just the tip of the iceberg. One will need to spend some time with this tool to really get into its abilities such as colour coding source node, information levels, and so much more!

Please check the Message Analyzer Blog for more information.

Note that an absence of System Centre and its components is deliberate. We find, at least at this time, that Failover Cluster Manager provides a far superior cluster management experience.

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Friday, 11 April 2014

Cluster-Aware Update Runs: How Long?

When one needs to provision a new Windows Server 2012 RTM/R2 cluster one of the time factors to keep in mind is the Cluster-Aware Update runs that will happen once the cluster has been brought up.

image

The above is a new Windows Server 2012 R2 four node cluster that we finished configuring last night.

The first CAU run started around 2300Hrs. As can be seen each node took about an hour to run through the process.

Even though we are not too far into the product release cycle for R2 one needs to keep in mind the update maintenance process if one is not using an up-to-date image for deployment. This would be especially true if deploying 2012 RTM or 2008 R2 clusters.

Cluster node configuration:

  • Intel Server System SR1695GPRX2AC with RMM
    • Intel Xeon Processor X3470, 32GB ECC, 120GB Intel SSD RAID 1, Dual Intel SAS HBAs

Some important Cluster update tracking links:

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Monday, 31 March 2014

Scale-Out File Server Cluster: Where Is the Witness Disk?

We were in a bit of a quandary as far as the how/what/where for a Witness Disk (quorum disk) for Scale-Out File Server (SOFS) cluster.

Our setup:

  • (2) Intel Server Systems R2208GZ4GC
    • Dual SAS HBAs, Dual Intel X540T2 10GbE
    • Scale-Out File Server Nodes
  • (2) Intel Storage Systems JBOD2224S2DP JBOD
    • Mix of 10K SAS, 15K SAS, and SSD SAS drives mirrored in each JBOD
  • (2) Intel Server Systems R2208GZ4GC
    • Dual Intel X540T2 10GbE
    • Hyper-V Nodes

In the case where we had a Hyper-V Failover Cluster using Direct Attached Storage via Promise VTrak E610sD dual SAS, Dell MD3220 dual SAS, or HP P2000 dual SAS we would set up a 1.5GB shared LUN for the quorum disk.

Well, in this case we don't have the logic in the storage to do that.

So, where do we put our Witness Disk for a SOFS based Storage Spaces cluster?

The only real clue we had was in Jose Barreto's blog post:

Specifically this line in the Storage Spaces setup:

New-VirtualDisk -FriendlyName Space1 -StoragePoolFriendlyName Pool1 -ResiliencySettingName Mirror –Size 1GB

An e-mail to Jose asking about that command, and also an e-mail to my fellow Cluster MVPs, came back with Jose confirming that the 1GB Virtual Disk (Storage Space) would be used for the Witness Disk.

We have two options when it comes to configuring the disk. Either we set it up as per Jose's blog post and verify in Failover Cluster Management (FCM) that the Witness Disk is the allocated 1GB or we stand up the cluster, configure the 1GB Storage Space, and then designate the Witness Disk in FCM.

Further reading:

Thanks to Jose Barreto and my fellow Cluster MVPs that answered the N00b questions! :)

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business

Friday, 28 March 2014

Flash Multiple Intel or LSI RAID and SAS HBAs Together

We ran into a bit of a question mark around updating the SAS HBAs (Host Bus Adapters) in our Scale-Out File Server nodes. The question being how do we update all of the SAS HBAs in the server without having to take things a part physically.

When looking at the HBA.NSH file that comes with the firmware update for EFI Shell we see the default command line in the batch:

  • sas2flash -f Intel\gb.fw -b mptsas2.rom -b x64sas2.rom

A quick search on the SAS2Flash utility and we found the LSI manual here:

In the guide we find what we need:

  • sas2flash -fwall Intel\gb.fw -b mptsas2.rom -b x64sas2.rom

Once we power cycled through POST and CTRL+C into the HBA BIOS we saw:

image

Happiness is not having to pull apart the server systems to update things one at a time!

Philip Elder
Microsoft Cluster MVP
MPECS Inc.
Co-Author: SBS 2008 Blueprint Book

Chef de partie in the SMBKitchen ASP Project
Find out more at
Third Tier: Enterprise Solutions for Small Business