I'd like to thank the patience of customers over this - as I said before, the main delay was the new fibre links between the London data centres we use, and we had planned to be at this stage late last year, before there were any capacity issues.
The good news is that the new switches are deployed (last Thursday). This took a bit longer than expected but was all done within the planned work window. Yesterday evening we were able to start using the new fibre links and relieve the congestion on our Talk Talk back-haul.
The next step is not quite what we originally expected. We were expecting to need more BT back-haul links, but the recent moves of many lines from BT to TT had meant we actually have enough capacity to both BT and TT for now.
Surprising the next bottleneck is the LNSs themselves. Only a few months ago we moved from 3 active LNSs to 4, and we have several more waiting to be plugged in. Sadly the original design linked the LNSs to BT back-haul, so we can't simply add more just yet - not without change.
The next step is therefore a change to the way we link to BT and TT back-haul, and that is planned for Sunday morning starting at 2am. This will mean that the number of BT links, and the number of TT links, and the number of LNSs will all become independent, and we can deploy new LNSs next week.
I hope that is one of the last major disruptive bits of work we have to do. Once we have that change, it will be possible to go back to our normal rolling over-night upgrades on LNSs when needed. It will also be possible to deploy new LNSs without major disruption.
After all of that we have to look carefully at our transit and peering - again, we have enough capacity for now, but it may need a bit of juggling to ensure no links are getting hot.
We can then take a bit of a step back and work out where we need more capacity next, and get it in place before needed.
The good news is that now the TT back-haul is not the issue, I am launching the Home::1 ADSL Terabyte service later today. Thank you for your patience.
2016-03-30
2016-03-26
Internet Connection Records, a small taste of the problems with #IPBill
However, there have been a small number of consequences which we have been working on. Obviously not show stoppers otherwise the planned work would have been reversed, but oddities.
One of them was that we were having difficulty getting SNMP from some of our LNSs, which meant some of our monitoring was unavailable. This had left us scratching our heads somewhat as the LNSs were not rebooted or reloaded or anything.
Then, another snag was that today one of our servers that does syslog started to run out of disk. Again, a puzzle. But this was easier to understand just by looking at the logs.
It turns out these are related. We have some debug logs from the LNSs related to setting up PPP sessions and allocation of IP addresses. These are kept for a couple of days to help resolve any connection problems.
One of the things logged is the IPv6 allocation, and this is logged by logging the DHCPv6 request/reply exchange from the customer router. Usually these either happen once after connection or maybe once an hour.
The problem, it seems, is rather odd. Some customers still use the Technicolor ADSL broadband routers that we used to sell from years ago. It seems many of these got upset in a rather odd way after the work on Thursday. We can see no logical reason for this, but they are now in a state where they are on-line and working, but generating approximately 1GB of uplink traffic a day, each, sending DHPCv6 requests! We were logging all of these. It seems the logging may actually have been so much load that it was impacting the SNMP responses.
The fix is rebooting the Technicolor routers, which, thankfully, we can do remotely.
But this gives me a slight insight in to the difficulty of collecting Internet Connection Records. Each of these DHCPv6 exchanges would be something that might well be logged as an ICR.
In practice, just trying to log this one type of packet we could not keep up - the log file was only 16GB (158 million entries) since 4am today. Looking at the traffic levels, that is a tiny fraction of the number of requests being sent by these routers. Our LNS logging system has built in limiting to try and avoid overloading things, and it was being pushed to the limits.
If we had to log every session (TCP/UDP/SCTP/IPSEC/ICMP, etc) there is just no way any of our existing kit could keep up. Of course it wasn't designed to! It was designed to shift packets quickly and provide Internet access to our customers, not snoop on anybody.
This also highlights the issue with any deliberate generation of ICRs by s/w on customer networks. It is easy with relatively low levels of traffic to cause a lot of ICRs to be created, if the #IPBill passes.
2016-03-25
CISCO and ARP?
The FireBrick has quite a good ARP handling subsystem, including exponential back-off, configurable ARP timeouts and so on. It has served us well, but we have recently encountered a slight problem talking to a CISCO Nexus switch.
So I did some tests - and would love to know if this is typical. Any CISCO experts reading this may be able to comment.
Testing using arping from linux, I could see that the CISCO would respond to only some of my ARP requests. Maybe one in five, but not very consistent. This is a tad odd, and may be down to some general ARP rate limiting perhaps.
On top of that, when it did respond, it did so after 2.99 seconds. This was very consistent - I had to use arping one ARP request at a time to confirm this.
I have to wonder what the hell it is doing! From a coding point of view, holding on to the ARP request or reply for that length of time is more work than just answering the ARP right away. I am at a loss as to what is going on.
For comparison, a FireBrick is timed by linux at 180us response and answered every ARP.
Anyway, it means I have had to tweak the way the ARP system renews ARPs to try a bit longer, otherwise every now and then the CISCO vanishes for a few seconds.
Oh, and yes, they still look like this with some arbitrary padding to min packet size for Ethernet.
P.S. It was CoPP, but we don't understand why it would delay ARPs 3s in that process.
So I did some tests - and would love to know if this is typical. Any CISCO experts reading this may be able to comment.
Testing using arping from linux, I could see that the CISCO would respond to only some of my ARP requests. Maybe one in five, but not very consistent. This is a tad odd, and may be down to some general ARP rate limiting perhaps.
On top of that, when it did respond, it did so after 2.99 seconds. This was very consistent - I had to use arping one ARP request at a time to confirm this.
I have to wonder what the hell it is doing! From a coding point of view, holding on to the ARP request or reply for that length of time is more work than just answering the ARP right away. I am at a loss as to what is going on.
For comparison, a FireBrick is timed by linux at 180us response and answered every ARP.
Anyway, it means I have had to tweak the way the ARP system renews ARPs to try a bit longer, otherwise every now and then the CISCO vanishes for a few seconds.
Oh, and yes, they still look like this with some arbitrary padding to min packet size for Ethernet.
09:40:20.688429 ARP, Request who-has 91.240.176.1 tell 91.240.176.254, length 46
0x0000: 0001 0800 0604 0001 0003 971d c009 5bf0 ..............[.
0x0010: b0fe 0000 0000 0000 5bf0 b001 474e 5520 ........[...GNU.
0x0020: 5465 7272 7950 7261 7463 6865 7474 TerryPratchett
P.S. It was CoPP, but we don't understand why it would delay ARPs 3s in that process.
2016-03-22
Public Bill Committee Written Evidence #IPBill
I am sorry that it has taken so long to get together a proper report on my findings when considering this Bill. It has been a lot of hard work, and I am very grateful to the assistance of many colleagues on this, including working through much of the Bill with me page by page on a Sunday!
There are, as before, a lot of issues.
My submission (PDF).
If you are thinking of making a submission, please do so ASAP. The first oral hearing is Thursday 24th and they have asked for evidence by Wednesday to allow time to consider it.
Please do feel free to quote me or copy to your MP. Ultimately they are the ones that vote on this.
I am happy to meet with MPs and Lords on this matter.
Oh, damn, once again, under this Bill your computer would now be logged as accessing a porn site, just because you read my blog. It would not log that you only accessed a benign favicon image, as that would be content, just that you accessed something on the site. Oops.
There are, as before, a lot of issues.
My submission (PDF).
If you are thinking of making a submission, please do so ASAP. The first oral hearing is Thursday 24th and they have asked for evidence by Wednesday to allow time to consider it.
Please do feel free to quote me or copy to your MP. Ultimately they are the ones that vote on this.
I am happy to meet with MPs and Lords on this matter.
2016-03-21
Signal #IPBill
Signal is a simple app for your phone - and you should install it and use it?
Why? well, for one simple reason it allows both iPhones and Android to message each other using data and not SMS or MMS. It also allows calls via data.
But the real reason is privacy - what you send and receive or say using signal is private.
It is free and literally took a matter of seconds to install and start using. It ties in to your phone number and contacts and just works.
But wait a second! Encryption is difficult because of validating keys. People have "key signing parties" for things like PGP email. How can you tell the person you are talking to is the person you think they are?
Well, signal actually makes that easy too - you can easily, when you meet someone, point your phone at their screen and it reads a fingerprint of their key and checks it matches. If ever it changes later the phone will tell you that there is a problem.
They make privacy simple, and the Investigatory Powers Bill has nothing in it to allow snooping on your texts and calls in the network, when using Signal. It does not outlaw using Signal either (would be hard to without outlawing https for access to banks too). As worded now it could try to order this non UK company to put in a back door, and they are pretty guaranteed to tell them to sod off. The source code is available and inspectable so even if they were compelled to comply it would be obvious.
You can even secure the message archive in the phone independently to any encryption the phone offers.
So, download, install and use Signal - why not?
As used by members of the House of Lords to protect their privacy...
Why? well, for one simple reason it allows both iPhones and Android to message each other using data and not SMS or MMS. It also allows calls via data.
But the real reason is privacy - what you send and receive or say using signal is private.
It is free and literally took a matter of seconds to install and start using. It ties in to your phone number and contacts and just works.
But wait a second! Encryption is difficult because of validating keys. People have "key signing parties" for things like PGP email. How can you tell the person you are talking to is the person you think they are?
Well, signal actually makes that easy too - you can easily, when you meet someone, point your phone at their screen and it reads a fingerprint of their key and checks it matches. If ever it changes later the phone will tell you that there is a problem.
They make privacy simple, and the Investigatory Powers Bill has nothing in it to allow snooping on your texts and calls in the network, when using Signal. It does not outlaw using Signal either (would be hard to without outlawing https for access to banks too). As worded now it could try to order this non UK company to put in a back door, and they are pretty guaranteed to tell them to sod off. The source code is available and inspectable so even if they were compelled to comply it would be obvious.
You can even secure the message archive in the phone independently to any encryption the phone offers.
So, download, install and use Signal - why not?
As used by members of the House of Lords to protect their privacy...
Replacing switches
The first step in upgrading our network is replacing some of the core switches with new, much faster and more powerful, switches.
Replacing switches is always fun!
For a start, they are in pairs to try and ensure continued operation of at least some of the network if one was to fail. Where possible devices are connected to both switches, and where we have pools of devices they are spread between the two. We actually have some new changes in the pipeline that will allow more of our equipment to actually use link aggregation over two switches for better redundancy even.
So, to move to a new switch, what do you do?
Well, first off, and surprisingly, you have to make space - you need the new switches basically next to the old ones in the rack. This may not be obvious, but if you are moving cables from one switch to another you need to make the move as short as possible. If not, then you have to re-route the cables or even get longer cables. So you have to shuffle stuff up/down to make space. Thankful that worked well. You also have to check cables are going to be able to move, and none are too short or snagged on anything.
Then, you make sure the new switch is the same config as the old. This is not simple as switch configuration is far from standard. There are VLANs and jumbo frames and all sorts to check very carefully. A lot of double checking is needed.
You also configure the old and new switch so that all of the VLANs can link between them. This means you can plug the new switches in to the old ones.
Then, on the day, you move one cable at a time. Ideally, shutting down operations of what you are moving cleanly to fall back to other devices, and then move the cable, check it, re-enable the functions, and check that. One by one very carefully. Done right you can move a lot of things with no impact on service at all - pairs of BGP servers can cleanly switch over, move, and switch back. Some things have disruption like LNSs which cause traffic to reconnect to other LNSs when shut down.
There can be (and were) problems! Basically the old switches had a head fit after moving many of the cables! This makes no sense, and meant power cycling the damn things. And, of course, moving cables back. It was not pretty.
We have tried this twice, and the second time we have Talk Talk suffer a major issue as well which complicated matters so even reverting the changes left us with all TT lines off line for a couple of hours.
So, this time, on Thursday, new approach, called "big bang". The same careful config, and checking, but not linking the old switch, just carefully but quickly moving every cable to the new switch and then spending time checking each one. It will cause more issues than the more usual step by step approach (when it works), but it is pretty predictable that it should actually work this time. However, there will be a clear time limit and move all the cables back if we cannot get everything working within that time, in the middle of the night.
Good luck to the ops team doing this work...
Replacing switches is always fun!
For a start, they are in pairs to try and ensure continued operation of at least some of the network if one was to fail. Where possible devices are connected to both switches, and where we have pools of devices they are spread between the two. We actually have some new changes in the pipeline that will allow more of our equipment to actually use link aggregation over two switches for better redundancy even.
So, to move to a new switch, what do you do?
Well, first off, and surprisingly, you have to make space - you need the new switches basically next to the old ones in the rack. This may not be obvious, but if you are moving cables from one switch to another you need to make the move as short as possible. If not, then you have to re-route the cables or even get longer cables. So you have to shuffle stuff up/down to make space. Thankful that worked well. You also have to check cables are going to be able to move, and none are too short or snagged on anything.
Then, you make sure the new switch is the same config as the old. This is not simple as switch configuration is far from standard. There are VLANs and jumbo frames and all sorts to check very carefully. A lot of double checking is needed.
You also configure the old and new switch so that all of the VLANs can link between them. This means you can plug the new switches in to the old ones.
Then, on the day, you move one cable at a time. Ideally, shutting down operations of what you are moving cleanly to fall back to other devices, and then move the cable, check it, re-enable the functions, and check that. One by one very carefully. Done right you can move a lot of things with no impact on service at all - pairs of BGP servers can cleanly switch over, move, and switch back. Some things have disruption like LNSs which cause traffic to reconnect to other LNSs when shut down.
There can be (and were) problems! Basically the old switches had a head fit after moving many of the cables! This makes no sense, and meant power cycling the damn things. And, of course, moving cables back. It was not pretty.
We have tried this twice, and the second time we have Talk Talk suffer a major issue as well which complicated matters so even reverting the changes left us with all TT lines off line for a couple of hours.
So, this time, on Thursday, new approach, called "big bang". The same careful config, and checking, but not linking the old switch, just carefully but quickly moving every cable to the new switch and then spending time checking each one. It will cause more issues than the more usual step by step approach (when it works), but it is pretty predictable that it should actually work this time. However, there will be a clear time limit and move all the cables back if we cannot get everything working within that time, in the middle of the night.
Good luck to the ops team doing this work...
2016-03-17
Call apparently from VERSO GROUP (UK) LIMITED - junk calls
Getting pissed off with junk calls today.
Two so far.
One was apparently from IDENTITY PROTECT LIMITED. Listen here.
The second, and much more amusing, was apparently from VERSO GROUP (UK) LIMITED. Listen here (posted with permission granted in the call recording itself). It seems that they were actually trying to sell me broadband. Wow!
I am not sure they were not the same person calling even, but that may be my being a tad racist... Listen and try and work it out yourself.
Update: Corrected first company name which was mistyped/linked.
Update: Clarified that second recording published with permission given within the call itself.
Update: For more details on why only "apparently" from Verso Group (UK) Ltd, see later blog post.
Two so far.
One was apparently from IDENTITY PROTECT LIMITED. Listen here.
The second, and much more amusing, was apparently from VERSO GROUP (UK) LIMITED. Listen here (posted with permission granted in the call recording itself). It seems that they were actually trying to sell me broadband. Wow!
I am not sure they were not the same person calling even, but that may be my being a tad racist... Listen and try and work it out yourself.
Update: Corrected first company name which was mistyped/linked.
Update: Clarified that second recording published with permission given within the call itself.
Update: For more details on why only "apparently" from Verso Group (UK) Ltd, see later blog post.
Subscribe to:
Posts (Atom)
Clocks
Some time geeks (should I say Time Lords) checked out my clocks. Seems they are impressed, saying sub microsecond. I have spent all day tryi...
-
Broadband services are a wonderful innovation of our time, using multiple frequency bands (hence the name) to carry signals over wires (us...
-
For many years I used a small stand-alone air-conditioning unit in my study (the box room in the house) and I even had a hole in the wall fo...
-
This is an appeal for (sensible) comments. I am working on revised A&A tariffs for broadband. For those that are not sure how they wor...
