2016-08-12

WoW IPv6

Having chased this as an ISP, we finally got someone in Blizzard who fixed their IPv6 issues.

It sounds like the front line people had no clue, closing faults as "resolved" even admitting it was not actually resolved.

Even though supposedly passing the issue on and the fact this is subject of several forum posts in US and EU, nothing happened.

Finally got someone in their ops, and he not only sorted the issue but set up the peering for us too. He did so in 3 hours of my email.

It feels like we could have sorted weeks ago if the right people had realised.

But good, once again IPv6 is working on Blizzard, and we have direct peering on IPv4 and IPv6 to Blizzard for A&A now. All the better for low latency on our services which have good low latency in the first place.

Why OFCOM's ruling on BT charging Talk Talk SFI and TRC charges does not help most ISPs

As reported by ispreview, Talk Talk made a complaint about the amount BT plc t/a Openreach charged for Special Faults Investigation (SFI) and Time Related Charges (TRC). They argued that the charges were not cost based and much of their complaints are upheld by OFCOM. There is now much ranting over back-dating and refunding, and so on.

As you know SFI charges are a major problems for most ISPs. The costs, of there order of £160+VAT per visit, are massively more than the monthly, or even annual, profit from selling a broadband line in most cases. Either ISPs get stung or end users get stung, even when the work done was fixing a broadband fault, something that we should not pay extra for!

To explain the problem you have to understand the layers of services provided. I'll explain for ADSL as this is where it is most relevant.

1. BT plc t/a Openreach sell wires in the ground, a metallic path service. They are not selling broadband.

2. BT plc t/a BT Wholesale buy the metallic path from BT plc t/a Openreach that connects a home or office to the exchange; their own equipment in the exchange (DSLAM); their own back-haul across the country; their own BRAS equipment; and routers and links to ISPs. They sell broadband. Talk Talk Business do the same as BT plc t/a BT Wholesale in this respect.

3. Companies like A&A buy broadband from BT plc t/a BT Wholesaled Talk Talk Business; have our own routers; DNS servers; and make use of transit and peering and equipment in data centres. We sell an Internet Access Service.

The complaint Talk Talk made is against BT plc t/a Openreach, and is only about the price. It is sensible for BT plc t/a Openreach to sell a service (SFI) that finds and fixes broadband faults as only their engineers can work on the network. It is sensible for BT plc t/a Openreach to charge for that, because it is over and above the metallic path they sell.

The issue we have, repeatedly, with BT plc t/a BT Wholesale and Talk Talk Business is that we have no interest in buying an SFI service. We already buy, and pay for, working, broadband and any work to fix that broadband is the responsibility of BT plc t/a BT Wholesale or Talk Talk Business. The fact that BT plc t/a BT Wholesale or Talk Talk Business have to pay BT plc t/a Openreach to find and fix broadband faults is not our problem, any more than the fact they may have to pay CISCO engineers to fix a BGP router in their network.

So our gripe has never been with BT plc t/a Openreach.

This ruling may make some SFI visits cheaper but as we should never pay for an SFI visit ever, and it is not a service we want to buy, all this ruling will do is make the amount we dispute every month slightly smaller. It won't solve anything useful for us. Sorry.

Even so, well done to Talk Talk on this - it will reduce their costs which is good news.

2016-08-10

Making network cables

More of a bit of Vlogging this one. I am working on my skills in this area, just for fun, and my office at home is a bit echoey. So I used a proper audio recorder this time instead of the one on the camera.

Final Cut Pro works well to allow me to select the camera 4k video, and the audio file and synchronise. I can then choose any mix of the two sets of audio tracks on the composite clip. In practice, just picking the audio from the recorder.

I also managed to take most of this using a 100mm macro lens, on fixed focus. Very hard to do this all by yourself, as you can see from some of the positioning. This was after a couple of takes (which is not like me). I had to set f/32 to get focus to be OK, as auto-focus could not track may hands well enough, hence slightly grainy, but I think it looks good.

I suspect my next challenge will be lighting. It would have helped with the graininess.

See what you think - it is just an educational video on Making cat5e network cables, that is all.

2016-08-05

iPhone Unifi DHCP issue

This is technical.

For a long time now my son James, and I, have been cursing Apple. We have iPhones, and iPads and all sorts, but we keep finding the WiFi not working in the house.

To explain the symptoms, in the morning I normally get up and have a bath and use my phone in the bath to check Facebook and twitter and so on, and every damn day I find my phone stops working in some way on the WiFi. Basically Facebook shows blank panes and not loading or things don't show new stuff. I have to go to WiFi settings and I see a 169 IP address (no response to DHCP standard address) as per the image on the right. The fix is WiFi off/on or airplane mode on/off. Sometimes I have to do this two or three times. It pisses me off.

I have done a lot of work on this. The set up is as follows :-
  • iPhones and iPads, latest code, James has beta code
  • Unifi APs, latest code, three of them, all same network and SSID
  • FireBrick doing DHCP
All of these are pretty solid systems, and should not screw up like this. I did loads on the FireBrick DHCP (seeing as I wrote it) trying everything I could think off - tweaking the TTL on responses, changing lease times, all sorts. Nothing helped.

I have pinned down that the problem happens on change of AP, when I get in the bath I am between APs and it moves from one to the other, and that is when it loses DHCP. To add to the fun, it has IPv6 working OK but not IPv4, so extra special.

Firstly, this shows a clear bug in the Apple code - there is "WiFi Assist" to handle poor wifi (and use mobile), but that does not kick in when you have working IPv6 and no reply on IPv4 DHCP. It knows we have no IPv4 on WiFi (hence 169 address). Maybe it should, at least, use mobile for IPv4 traffic!

But my packet dumps suggest no attempt to get DHCP on these cases. We see IPv6 working, but IPv4 packets to the phone, and ARPs to the phone do not work, and there are no DHCP requests.

James spotted the broken "WiFi Assist" and we tried turning that off, but sadly no joy. It is not 100% reproducible, so hard to be sure we have found a fix, sadly. But this did not work.

So what next? Is it an Apple bug, or a Unifi bug, or even a FireBrick bug? So we tried some more.

In some degree of desperation we set all three APs to different SSIDs. This is to see if that works, but the problem is the phone sticks to the wrong SSID even with really low signal as we move around the house. Yes, when/if it changes AP it gets a new IP, but it sits on the crappy signal for ages. It does not switch until it totally loses signal. Bugger.

So what next. Well, I have made the cardinal sin of changing two things at once, and this seems to be working. If I can, I'll update and confirm the diagnosis with one change later.
  • I set up the SSID we are using only on 5GHz not 2GHz
  • I changed the Unifi to consider all three APs to be separate Wireless Networks that happen to be on same LAN (VLAN not set) and happen to be same SSID, rather than one network on all three APs
I have tried many times to break it now, and when moving from one AP to another the phone quickly re-associates with the nearer AP and does not lose its IP addressing at all. It seems to be working! The real test will be the bath, tomorrow.

So, jury still out on if an iPhone issue of a Unifi issue. But we have working WiFi at last.

When we come to reporting this to one or other of them, we will not get far, I am sure.

Update: In spite of all my testing, it just went wrong again, arrrrg! What is worse is that it just sits there with a 169 address not even trying.

Update: I tried changing all the APs to actually be on the same channel, and setting a min rssi. That seemed to help in that nothing broke this morning. We'll see how it goes for a few days. Several of the comments suggest this is definitely an Apple issue but at this stage I am trying to work around it.

Update: The 5GHz has helped, but sadly, today, it failed again, so trying same SSID and channels did not help - next step is try different SSID and set min-RSSI

Update: Using different SSIDs does fix this, and setting min RSSI means they do switch. Not ideal. I tried working our Enterprise WPA but not worked out the RADIUS responses I need yet (If anyone has a pcap that would be helpful). One other simple thing to try would be fixed IP rather than DHCP.

Captive portals (apple)

You will have noticed are those rather annoying captive portals you get on hotspot / WiFi. They are frustrating in a lot of ways, and even "free" WiFi can often mean completing some details or even just pressing an "I agree to terms" button. As recently explained by arstechnica, running a WiFi does not usually mean any sort of filtering, logging, or terms and conditions are actually needed. Indeed, in some cases, having T&Cs could form a contract where one is not needed and that can result in some obligations that you don't want. It really bugs me that you can rarely use a free WiFi with a device that has no browser, like my camera, as no way to get passed the splash screen.

Much as I hate these damn things, what is quite nice is the way Apple devices pop up the splash screen when you select a WiFi, and if you don't login/accept, then it does not use that WiFi. This is Apple being a bit clever and actually the way they do it can be used sensibly.

How do they work?

The way a portal, or pay wall, or whatever you want to call them, usually works is by not providing a working Internet connection at all! They divert "web pages" to a splash screen for login/etc. How this divert works can vary, it could be block all traffic and override DNS, it could be redirect port 80. I have seen some redirect port 443 which creates a nasty security warning and really is not a good idea. Changing DNS can result in nasty caching effects.

Once you complete the process the diverts are removed and normal Internet access is possible.

What do Apple do on devices?

What apple do when you select the WiFi is make a simple HTTP GET request (not https for obvious reasons) to http://captive.apple.com/hotspot-detect.html

The response is a simple http page with Success in the title and contents. If the device sees that then it assumes it has working Internet access, and uses the WiFi.

If not, then it displays whatever page it gets instead. This works well with these typical arrangements that divert all web pages.

These "test" requests come from CaptiveNetworkSupport-325.10.1 wiser user agent, but requests to display the page and subsequent pages to complete the login are from the normal Mozilla/5.0 (iPhone; CPU iPhone OS 9_3_3 like Mac OS X) AppleWebKit/601.1.46 (KHTML, like Gecko) Mobile/13G34 user agent.

Sadly sending cookies does not seem to work, so there is no real way to tell the initial test from any further test, so the server has to have some state. Normally the server will have state, knowing the user is allowed access or not (yet), and when allowed the test will go to the real captive.apple.com and be served correctly.

What if you want a simple splash page?

What I wanted to do was make a simple splash page, no terms and conditions, just saying who is providing the free WiFi. The idea is to do this for an free WiFi in a cafe. I want the WiFi to work as well as possible, with IPv6, and to work on devices without browsers if possible. But for iPhone users at least, a popup splash page would be nice.

So, first thing, divert DNS for captive.apple.com to my own server. The good news is that this can be a permanent fix in the DNS server used for this connection, it does not need to know the user is allowed or not, an we are not setting any general blocking of anything. It is a single DNS entry override with IP address. By the way, this used to be a page on http://apple.com/ which would have meant redirecting the whole apple.com domain in DNS, thankfully Apple have changed this to a specific subdomain now.

The server then needs to serve something for hotspot-detect.html to provide the splash page. However, after each page the phone re-checks hotspot-detect.html to see if now allowed or not, so I have to set some state so that the second and subsequent requests (in a short time frame) serve the expected Success page.

You don't see the MAC unless you have something local, or you can ask the remote device to use ARP or ND to find out. All you see is requesting IP address. I also have concerns that the IPv6 is a privacy address and likely to change. It may be possible to only serve an IPv4 address at DNS level to avoid that issue but that is messy.

My solution was simply to mark that IP as allowed for a short period so that only the first request to hotspot-detect.html would redirect to the splash page.

The result is the phone pops up with the splash page and is then immediately happy that it is now allowed and shows "Done" on the screen. This is exactly what I wanted.

Other devices just make use of Internet with no splash page, and there are no restrictions, which is also what I wanted.

Next step - see if Android phones do anything similar.

2016-08-03

Hard sell from BT

It has been a long day, and part of that has been arguing with BT, with Shaun and Alex's help.

We have a very understanding customer with a mostly working service and we are giving him a discount for indulging us on this point. But it has allowed us to push BT on the ongoing issue of SFI engineers yet again. I hope that we will eventually get a straight answer and I can update this post.

Simple story - a phone line working OK for phone calls but dropping broadband frequently. Already engineers have established it is an issue with the drop wire and need to change it, and fit a new anchor to the building, and so on. I.e. the fault has been properly investigated by BT and remedial work proposed. This is all good.

However, the fault has stalled, and last week we escalated to High Level Escalations. On Monday we are told that BT plc t/a BT Wholesale cannot even talk to BT plc t/a Openreach unless we boot an SFI engineer!

At this point it is worth explaining that BT plc t/a BT Wholesale define an "SFI Engineer" as an optional service we can request that will send a man to check the line meets the "metallic path specification" (SIN349). As far as we know the line meets SIN349, and we have no reason to order such an optional service in this case (or indeed, ever!).

We have endured THREE DAYS now of BT trying their damnedest to sell us this optional service in order to progress the repair of a fault on a broadband service. Note that broadband is not measured against the "metallic path specification" anyway, we don't buy a "metallic path", we buy "broadband". So it is a pointless service, and one they know we will be charged for as the line meets SIN349 (they charge for SFI if the line meets the spec).

The fact it has taken three days to make no progress is why we need an understanding customer. We could have booked the optional extra service of an SFI and the fault would have progressed, and then we would have to dispute the charges later. But we want this issue resolved so we don't have these issues and disputes in future.

At this stage I have now asked, at least half a dozen times, of account manager, and High Level Escalations, and other BT staff :-
Obviously you have a process to get broadband faults fixed which does not involve us ordering an optional extra service. That stands to reason, as otherwise you would not be able to meet your contractual obligation to fix broadband faults.

So please, just tell us the process we have to follow to do that, and we can get this fault fixed.

And...

At this point we have simply asked that you tell us the correct process by which we get BT plc t/a BT Wholesale to fix a broadband fault without us ordering an optional extra service.

It is surely a really simple question you can just answer for us right away. Once we know, we can follow it and this fault can be resolved.

Obviously you must have such a process else how would BT meet its contractual obligations to rectify broadband faults.

So please, just tell us..
And variations on that...

BT should be able to just answer that simple question. For some reason they just pass the buck and do not answer. Oh, and magically, 6pm today, they finally think they can progress this fault without us booking an SFI engineer. But clearly a process that involves three days of arguing and BT trying to sell an optional service to us, cannot be the proper official process, so we still await the answer.

I'll keep you posted :-)

P.S. ispreview have picked this up (thanks) and it is worth clarifying...

The BT plc t/a Openreach side is not that daft – they define SFI differently, and they sell that service as a means to help address broadband faults. They charge because (at least for ADSL) Openreach don’t sell broadband, they sell metallic paths.

The issue here is actually BT plc t/a BT Wholesale, not BT plc t/a Openreach. BT plc t/a BT Wholesale redefine SFI as a service simply to check the line to SIN349, and not a service to fix broadband faults. After all, we would not want to buy a service to fix broadband faults as that is already part of the broadband service we already pay for. That redefinition means we would never want to buy that SFI service. However, logically, BT plc t/a BT Wholesale would want to buy the SFI service from BT plc t/a Openreach to help fix broadband faults (as they are the only people that can work on the line), and as fixing broadband faults is part of the broadband service, they should not charge us when they pay for that service. I hope that make some sense.

Update: So far (5th Aug) the best I have is: "BT will take steps to repair a broadband line that is not operating within its contracted specification or fault threshold. The exact nature of the resolution will be determined as part of the diagnostic journey and will differ depending on which component or product is affected. The process for contacting BT Wholesale and raising a fault is contained within the customer service plan." which does not really answer the question.

2016-08-02

Up to 80Mb/s

Once again there are calls to change the way "broadband" is advertised [ispreview], describing it as misleading.

I have said this before, but once again I'll try and explain why the proposals make no sense. It was rather telling that 10% of people surveyed were happy with the current system. The current system involves lying about the technology, claiming it can only achieve a maximum speed that is a speed 10% of people could get, e.g. 76Mb/s. instead of 80Mb/s That means 10% of people will get more than that maximum and so have been lied to about it being the maximum. It means that if people feel misled by "up to 80Mb/s" now you have 90% of those people still feeling misled by "up to 76Mb/s". The problem is not solved, you simply have 10% fewer complaints about being "misled".

I do wonder if a simple solution is changing "up to" to "not more than". Would people still feel misled? I bet they do somehow.

Personally I think people are not clear on the way the technology works, and that is a hard one to solve. Let's consider two technologies here that are pretty simple, and for which people may have a choice. One is ADSL to the exchange, using ADSL2+ modem standards, and this allows up to 24Mb/s sync rate on the line. The other is VDSL to the street cabinet, and this allows quite high speeds but is sold by Openreach in a package with an 80Mb/s sync limit, so up to 80Mb/s sync rate on the line.

If you have a choice of "ADSL or VDSL" they are technical terms and the average consumer has no clue.

If you have a choice of "up to 24Mb/s" or "up to 80Mb/s" you can guess which is better and have some rough idea of the scale of "betterness" that may apply. Note that actually, it is possible, for long lines to get better speed on ADSL than VDSL in a few cases, but that should be clear when actually ordering. The key here is that the different technologies can be simply compared. Indeed, a 330Mb/s FTTP service is clearly "better" still, where as 8Mb/s ADSL1 is "worse", and so on.

So for comparing the technologies, the "headline" speed in the "up to" claims is a perfectly sensible comparison for most consumers.

The problem then comes when people actually buy a service and find it is not, for example, 80Mb/s.

The guidelines are very clear on this and I believe most ISPs follow them. When someone orders, they are told an estimate of the speed they are likely to get where they are. If you order from A&A, and are told 25-30Mb/s you cannot really complain that you have been "misled" with "up to 80Mb/s" when you actually get a speed within that estimated range, e.g. 28Mb/s.

There are issues as the lines may not get the sync speed within the ranges we quote, that can happen. These are estimates are based on BT data which we cannot control. Nobody seems to be complaining that these estimates are wildly wrong though, which is good. That probably means we have a system that works.

So for a start I really do not understand why we are having these cries of people feeling misled in the first place. Who are these people buying a service with a specific speed range estimate and then getting upset that the speed is within that estimate but not the "headline" speed advertised?

Anyway, there are a few other problems. For a start I think people really struggle with the idea of speeds of broadband. E.g. you don't use the full speed all the time - you use a certain amount if streaming a video, or playing a game, you may use the maximum when downloading something or maybe not. I think people see that, but I think they expect that if the service is "up to 80Mb/s" that they can push it that far if and when they need, at least some of the time. People seem not to realise that the limit will be constrained by their line and location and that is not something they can generally change in anyway, and neither can the ISP.

People also don't seem to appreciate that there are a relatively small set of technologies available, and in many cases actually the same wires and modems and equipment and backhaul between different ISPs - so if one ISP sells Openeach FTTC GEA as "up to 76Mb/s" and another as "up to 74Mb/s" they will actually be buying the same thing from either, and it might be only 50Mb/s, so the comparison of arbitrary per-ISP 90th percentile is not actually helping people decide.

The other problem is people do not know what the speed actually means, if we say 25-30Mb/s and the line gets 28Mb/s, then great, but what does that 28Mb/s mean? It means that the modem to modem sync rate provides 28Mb/s throughput (downlink). It does not mean you can download a file from google at 28Mb/s. Why? Well, a lot of reasons...
  • Your computer will have limitations on what it can do. To be fair, most modern computers are very fast and can easily exceed the speed of you line, but that will not always be the case as we get faster and faster technologies.
  • Your wiring and network in your home, and especially (if you are using it) your WiFi. These are limiting factors, and WiFi can be a big one.
  • The fact that your home network is shared (contended) with other people in the house, or your neighbour using your WiFi - this all creates a risk that things slow down due to other people's actions, even before we look outside your house!
  • The modem itself, probably provided by the ISP, should be able to handle the speeds. Depending on the equipment, that is not always the case though.
  • The modem to modem link, the sync speed of the line. That is a limiting factor. That is what the ISP is selling in terms of "speed", and that alone.
  • The link from cabinet to exchange for VDSL is shared (contended), so usage by other people own the same cabinet (regardless of ISP) can slow you down.
  • The link from exchange to BRAS is shared (contended), so usage by other people in your part of the country could slow you down
  • The link from the BRAS to the ISP is shared (contended), so usage by other people in the whole country could slow you down
  • The link in to the ISP themselves is shared (contended), and so other usage by the ISP's customers could slow you down. That is, in part in the ISPs control (see below).
  • The ISP network is, of course, shared, as are their links to peers and transit, so other usage by the ISP customers could slow you down. That is, in part in the ISPs control (see below).
  • The transit networks and peering points are shared (contended), so other internet users in the world could slow you down.
  • The link to the web site or server you are accessing is shared (contended), so other users of that service could slow you down.
  • The servers themselves will only be able to handle so much traffic, so other users of each server could slow you down.
  • The servers could impose specific rate limiting on what they will send, and that is their choice.
There are very few parts of this equation that the ISP actually controls. They control their links to the broadband network (and some have links they control right out to BRASs and even exchanges). They control their network and links to the transit and peering networks. They control their choice of backhaul and transit and peering networks, but there is often quite little choice.

However, these links will always be shared. To make them un-contended so that they can never fill up would mean costs rising by factors of many hundreds. It would make the broadband service ridiculously expensive, and would not eliminate all of the other aspects where the network is shared. There simply is no point. In practice ISPs do vary, and have different policies, but these are not something addressed by the ASA in any way or published by ISPs. At A&A we aim to keep links to carriers un-congested, i.e. we have enough capacity, but this is no guarantee that usage won't suddenly exceed it on occasion. Some ISPs run their networks so they hit limits when busy every day.

Ultimately what you are buying is a access (at a modem to modem access speed) to a shared global network, most of which is totally outside your ISPs control.

P.S. A&A have announced that Home::1 BT back haul FTTC will be one price now, dropping the extra £5 charge for the 80/20 option (from next bill). This brings the pricing in line with the TT back haul FTTC products. Yes, this means the same price whether the sync is 8M or 80M.

Clocks

Some time geeks (should I say Time Lords) checked out my clocks. Seems they are impressed, saying sub microsecond. I have spent all day tryi...